On Built This Week, AI infrastructure keeps coming back to one question: what does it cost to run this stuff at scale? Mitesh Agrawal of Positron AI builds inference chips around memory. Ben Lerner of Espresso AI uses machine learning to cut data warehouse compute. Karim Malhas of Breeze and Emanuele Melis of AI One build the layers enterprises use to run agents, and Amar Goel of Bito told us how hard it is to get model capacity at billions of tokens a day.
Everybody's competing to get the best agent, the best latency, the best all of these things, but people are forgetting. Once you get this, how are you going to run it?
— Karim Malhas, Episode 42
How companies we talked to use AI infrastructure
Positron AI: inference silicon built around memory
Mitesh says Positron is attacking the memory wall: both memory bandwidth and how much memory sits next to each chip. Its cards hold up to two terabytes, which matters when a 10-trillion-parameter model needs about five terabytes just for its weights. Positron targets the workloads where that pays off, like very large models, long contexts and image or video generation. (Episode 37)
His pitch is a "slider bar": pay more for very fast tokens per user, or run the biggest models with lower interactivity and a fixed budget.
Espresso AI: machine learning that right-sizes compute
Ben describes Espresso as "Kubernetes but powered by ML" for Snowflake and Databricks SQL. It gets an account inside the warehouse to resize clusters, and runs a proxy on the customer side that routes each query in real time. Its models read the SQL as it arrives and predict what adding one more query to a busy cluster will do. GPU inference is next on his list. (Episode 27)
Breeze: the infrastructure under voice agents
Breeze is provider-agnostic: you pick the speech-to-text (Deepgram in the demo), the model and the voice (ElevenLabs). Karim says Breeze has nailed the infrastructure part, from SIP trunks to latency and concurrency, and is now building the operational layer for the COOs who have to run these agents. One example is phone groups: one number making 10,000 calls a day gets blocked, so Breeze routes across many numbers and picks a local one. (Episode 42)
AI One: a context layer instead of a data migration
AI One's Context One sits between a company's systems and gives agents accurate, governed context without a big data transformation program. Every workflow starts with an "autonomy slider" at zero, where the agent only drafts and researches. Emanuele's favorite case was a regulated investigation where agents check the transactional database, documents and logs and produce a report in a couple of hours instead of a week. (Episode 46)
Bito: running into capacity limits
In the news segment on Amazon's Trainium 3 chip, Amar said Bito uses billions of tokens a day and model providers still tell it they lack capacity. His read: Nvidia still dominates training, while others are catching up on inference. (Episode 23)
What we built
- AI inference cost simulator (GPU vs Positron): Sam's build for Episode 37. It scales a chat workload on GPUs until margins collapse, switches to image generation where costs blow up, then reruns it on Positron. Mitesh said it was very close to Positron's own internal comparisons and asked to use it in their demos. He also asked where the numbers came from and stressed that Positron's edge only applies to certain workloads.
- Espresso AI data warehouse waste quiz: Sam's questionnaire for Episode 27 on SELECT * habits, dashboard refreshes and cost ownership, ending in a "CFO BPM" score and up to $8,000 a month in possible savings. Ben liked it but said it describes manual FinOps, "exactly the opposite" of Espresso's automatic approach.
What's still hard
- Supply. Mitesh said Positron avoids the HBM memory Nvidia uses because a startup can't get it, and uses commodity memory instead, even though its contract price rose two to two and a half times. Jordan said we have tried to buy GPUs online and they are hard to find.
- Energy. Mitesh said inference is limited by how much energy is available. Customers told him that more tokens per kilowatt-hour, and fitting into air-cooled data centers GPUs can't use, are critical.
- The math. Amar walked through the IBM CEO's point that $8 trillion of capex needs about $800 billion of profit a year just to cover interest, far beyond today's AI revenue.
- Latency against control. Every hand-off between voice agents adds delay, so Karim uses two or three agents at most, and checks codes like OTPs in deterministic code rather than trusting the model.
- Cost visibility. Jordan burned through tokens on a new model's top effort setting and said no tool tells you a command will cost $100 before it runs.
- AI-written infrastructure code. Ben still writes much of Espresso's backend by hand.
But, like, I can't just, you know, prompt at something and then throw it out into production and expect to run, like, tens of millions of dollars of Snowflake workloads through it. Like, it's it's just not that good yet.
— Ben Lerner, Episode 27
Episodes on AI infrastructure
- Episode 23 with Amar Goel of Bito: chip competition, GPU scarcity and capex math (December 2025).
- Episode 27 with Ben Lerner of Espresso AI: ML for warehouse compute (January 2026).
- Episode 37 with Mitesh Agrawal of Positron AI: inference costs and the memory wall (April 2026).
- Episode 42 with Karim Malhas of Breeze: running voice agents at scale (May 2026).
- Episode 46 with Emanuele Melis of AI One: context control for enterprise agents (June 2026).
FAQ
Why is AI inference so expensive?
Positron CEO Mitesh Agrawal points to the memory wall: very large models, long contexts and image or video generation need far more memory than a GPU holds, so the work gets spread across many chips. He adds that energy and chip supply are real-world limits on top of chip efficiency.
What does AI infrastructure include?
Mitesh describes three stacks under AI: energy, models and data, and silicon, networking and compute. Our guests also built software layers on top, like Espresso AI's compute optimizer, Breeze's voice agent platform and AI One's context layer.
Can AI optimize cloud data warehouse costs?
Espresso AI CEO Ben Lerner says yes for Snowflake and Databricks SQL: its models read incoming SQL, compare it with earlier runs, resize warehouses and route queries through a proxy, without the customer changing code.
Is AI good enough to write backend infrastructure code?
Not yet, according to Ben Lerner in January 2026. He uses Codex for dashboards and throwaway scripts, but would not ship prompted code that runs tens of millions of dollars of Snowflake workloads.