Built This Week/Guests

Guests

Mitesh Agrawal

Mitesh Agrawal, CEO of Positron AI and Lambda co-founder, explains why inference is a memory problem and how Positron's chips with up to 2 TB of memory each go after it.

Mitesh Agrawal is the CEO of Positron AI, which designs and builds silicon for AI inference. Positron's focus is the memory wall: how fast a chip can use its memory and how much memory it has. Mitesh was co-founder and COO of Lambda before joining Positron in January 2025. He joined us in Episode 37 (April 3, 2026).

And Positron is making the world's first terabyte plus memory chip, so having capacity of memory up to two terabytes.

— Mitesh Agrawal, Episode 37

What Positron AI does

Mitesh explained the problem with numbers. A 10-trillion-parameter model needs about five terabytes just to store its weights, even quantized. A GPU has a few hundred gigabytes, so the model gets spread across many GPUs. A Positron system has over nine terabytes of attached memory, so it fits in one box.

Positron's first-generation product proved out memory bandwidth utilization (MBU). The second chip, which he said tapes out later in 2026 and comes out in mid-2027, targets maximum MBU and maximum capacity. He added that Positron is one of the few early-stage AI silicon companies with tens of millions of dollars in revenue.

Key ideas from the conversation

  • Be honest about where you win (10:21). Mitesh called NVIDIA one of the world's smartest companies. Positron targets specific workloads (very large models, long contexts, image and video generation) where he says it can deliver 2x to 4x.
  • Code generation is mostly cache (11:30). About 90% of tokens in code generation are caching, which rewards lots of memory for hot caching.
  • Energy is the real constraint (09:03). Customers care that Positron can go into air-cooled data centers and produce more tokens for the same watts.
  • Cheaper tokens mean more tokens (12:55). Better compaction won't shrink demand. People will just upload a whole repo, then best practices, then more, especially with many agents feeding each other.
  • Commodity memory over HBM (14:17). HBM is the best memory, but startups can't get it. Positron uses the memory in phones and laptops, which has gone up two to two and a half times in contract pricing but is actually available.
  • CPUs are back (24:07). Positron is one of the first accelerators integrated with Arm's first CPU of its own. He said Intel had raised CPU prices twice in six months and still sells out.
What we wanna create and what we are creating is we wanna give the slider bar to the end user.

— Mitesh Agrawal, Episode 37

That slider bar, at 17:08, is the trade-off between speed and cost. If you want thousands of tokens per second on a gpt-oss model, you pay for more chips. If you run hundreds of agents on the largest model with a fixed budget, you trade interactivity for cost. Mitesh said most inference ASICs only offer the fast, expensive end.

What we built for the episode

Sam built an AI inference cost simulator that explains the inference problem, then runs a workload on GPUs as scale grows. Margins collapse, image generation sets off a red alert, and switching to Positron lowers compute load, latency and cost per million tokens. Sam called it "pretty juvenile." Mitesh disagreed: he said it was very close to Positron's own internal comparisons and asked to use it in their demos, while asking where the data came from. For more on this space, see AI in semiconductors and AI infrastructure.

FAQ

What does Positron AI do?

Positron AI designs and builds silicon for AI inference. CEO Mitesh Agrawal says it focuses on the memory wall: maximizing memory bandwidth utilization and putting up to two terabytes of memory on each chip.

Who is the CEO of Positron AI?

Mitesh Agrawal became CEO of Positron AI in January 2025. Before that he was co-founder and COO of Lambda.

Why doesn't Positron AI use HBM?

Mitesh says HBM is the best memory, but startups can't get allocation and it limits capacity. Positron uses commodity memory, the kind in phones and laptops, and uses its architecture to get closer to HBM on performance per dollar and per watt.

Where does Positron AI beat GPUs?

Mitesh was careful to say not everywhere. He pointed to very large models, very long contexts, and image or video generation, where he says Positron can deliver two to four times the performance of the latest NVIDIA chips.