Built This Week/Guests

Guests

Ben Lerner

Ben Lerner co-founded Espresso AI, which uses machine learning to cut Snowflake and Databricks SQL bills by resizing warehouses and routing queries in real time, with no code changes.

Ben Lerner is a co-founder and the CEO of Espresso AI, which uses machine learning to optimize data warehouse compute. Today it works with Snowflake and Databricks SQL, cutting bills without customers changing their code. Before Espresso, Ben worked on natural language processing at Google, on teams and code that later turned into Gemini. He joined us in Episode 27 (January 16, 2026).

What Espresso AI does

The goal of Espresso is to use machine learning to optimize compute. So we use ML models to rewrite code and make it faster.

— Ben Lerner, Episode 27

Espresso also shifts workloads between clusters and makes sure customers run the right machines at the right time. Ben's shorthand for technical folks: Kubernetes, but powered by ML.

It hooks in two ways (08:53). An account inside Snowflake or Databricks lets Espresso watch from a bird's-eye view and make warehouses larger or smaller. A proxy on the customer side then routes queries in real time, like air traffic control, so a workload that ran on 10 clusters might run on seven. It does not go through the Snowflake or Databricks marketplaces, which are not built for rerouting all of a customer's traffic.

Key ideas from the conversation

  • Transformers can read workloads, not just write code (12:08). Espresso's models read SQL as it arrives and compare it with previous runs. They predict whether an eighth query on a busy machine will queue, spin up capacity or slow the other seven.
  • Not doing work beats doing work (04:28). FinOps discipline helps, and at scale you need it. Ben's pitch is a button that cuts the bill so an overworked team of two data engineers doesn't have to.
  • GPU inference is next (10:04). Any large compute where you don't know your needs ahead of time can benefit, including packing uncorrelated workloads onto the same machine.
  • AI is not ready for serious backend work (13:55). Codex writes a lot of Espresso's dashboards and throwaway pandas scripts, and Ben hasn't visited Stack Overflow in over a year. He still writes much of the core code by hand.
  • 10 to 30% more productive (16:00). Espresso pays far less for Claude Code than a third of an engineer per engineer. They have no dedicated front-end engineer; those tasks go into the on-call rotation.
But, like, I can't just, you know, prompt at something and then throw it out into production and expect to run, like, tens of millions of dollars of Snowflake workloads through it. Like, it's just not that good yet.

— Ben Lerner, Episode 27

He expects design to become more of a bottleneck as front-end engineering gets easier. Customers saving millions want a product that looks better than one built without a designer on staff.

What we built for the episode

Sam built an Espresso AI data warehouse waste quiz. It asks a data team about SELECT * in production, their most expensive queries, dashboard refresh rates and who owns warehouse cost. It then returns a "CFO BPM" score and an estimated saving of up to $8,000 a month.

Ben called it a super fun quiz, then pointed out the irony: it describes FinOps practices, which is the opposite of what Espresso does. Espresso is not opinionated about what you run or why; it moves things around in the background.

On Gemini and Siri

In the news segment on Gemini powering Siri, Ben said the models keep pace with each other and the upgrade for Siri is the exciting part. He also shared that Daniel Gross is an Espresso investor. When GPUs were impossible to get in 2023, the team bought GPUs off Amazon and ran them in their office. For more on this space, see AI in AI infrastructure and AI in developer tools, plus our Claude Code vs OpenAI Codex comparison.

FAQ

Who is Ben Lerner?

Ben Lerner is a co-founder and the CEO of Espresso AI. He previously worked on natural language processing at Google, on teams and code that later fed into Gemini, and started Espresso about six months after ChatGPT came out.

What does Espresso AI do?

Espresso AI uses machine learning to optimize compute. Today it works with Snowflake and Databricks SQL, resizing warehouses and routing queries through a customer-side proxy so teams use less compute without changing their code.

How does Espresso AI reduce Snowflake costs?

It connects to the customer's Snowflake or Databricks account to resize warehouses, and runs a proxy that routes queries in real time. Its models read incoming SQL and compare it with previous runs to decide how much compute is needed each second.

Does Ben Lerner use AI to write code?

Yes, for dashboards, throwaway scripts and API glue, often via Codex. But he still writes a lot of code by hand and says models are not yet good enough for serious backend infrastructure.