Ben Lerner is a co-founder and the CEO of Espresso AI, which uses machine learning to optimize data warehouse compute. Today it works with Snowflake and Databricks SQL, cutting bills without customers changing their code. Before Espresso, Ben worked on natural language processing at Google, on teams and code that later turned into Gemini. He joined us in Episode 27 (January 16, 2026).
The goal of Espresso is to use machine learning to optimize compute. So we use ML models to rewrite code and make it faster.
— Ben Lerner, Episode 27
Espresso also shifts workloads between clusters and makes sure customers run the right machines at the right time. Ben's shorthand for technical folks: Kubernetes, but powered by ML.
It hooks in two ways (08:53). An account inside Snowflake or Databricks lets Espresso watch from a bird's-eye view and make warehouses larger or smaller. A proxy on the customer side then routes queries in real time, like air traffic control, so a workload that ran on 10 clusters might run on seven. It does not go through the Snowflake or Databricks marketplaces, which are not built for rerouting all of a customer's traffic.
But, like, I can't just, you know, prompt at something and then throw it out into production and expect to run, like, tens of millions of dollars of Snowflake workloads through it. Like, it's just not that good yet.
— Ben Lerner, Episode 27
He expects design to become more of a bottleneck as front-end engineering gets easier. Customers saving millions want a product that looks better than one built without a designer on staff.
Sam built an Espresso AI data warehouse waste quiz. It asks a data team about SELECT * in production, their most expensive queries, dashboard refresh rates and who owns warehouse cost. It then returns a "CFO BPM" score and an estimated saving of up to $8,000 a month.
Ben called it a super fun quiz, then pointed out the irony: it describes FinOps practices, which is the opposite of what Espresso does. Espresso is not opinionated about what you run or why; it moves things around in the background.
In the news segment on Gemini powering Siri, Ben said the models keep pace with each other and the upgrade for Siri is the exciting part. He also shared that Daniel Gross is an Espresso investor. When GPUs were impossible to get in 2023, the team bought GPUs off Amazon and ran them in their office. For more on this space, see AI in AI infrastructure and AI in developer tools, plus our Claude Code vs OpenAI Codex comparison.
Ben Lerner is a co-founder and the CEO of Espresso AI. He previously worked on natural language processing at Google, on teams and code that later fed into Gemini, and started Espresso about six months after ChatGPT came out.
Espresso AI uses machine learning to optimize compute. Today it works with Snowflake and Databricks SQL, resizing warehouses and routing queries through a customer-side proxy so teams use less compute without changing their code.
It connects to the customer's Snowflake or Databricks account to resize warehouses, and runs a proxy that routes queries in real time. Its models read incoming SQL and compare it with previous runs to decide how much compute is needed each second.
Yes, for dashboards, throwaway scripts and API glue, often via Codex. But he still writes a lot of code by hand and says models are not yet good enough for serious backend infrastructure.