For Episode 37 (April 3, 2026), Sam built a short explainer and simulation of AI inference costs: run a workload on traditional GPUs, crank up the scale until the margins break, then switch to Positron's chips. Our guest, Positron CEO Mitesh Agrawal, reviewed it live. Sam didn't say how long it took or what he built it with, and he opened by warning it might be "pretty juvenile."
We wanted viewers on the same page about inference economics before talking to someone who designs inference silicon. The simulator's opening slide sets up the whole argument:
While training models get all the headlines, inference is where 90% of the real world AI costs happen.
— Sam Nadler, Episode 37
From there the slides argue that GPUs were built for graphics, not text generation, so you pay for idle compute, and that AI is limited less by intelligence than by how efficiently we can run it.
The episode didn't cover the tools behind the simulator, so we won't guess. What it showed on screen:
Better than Sam expected. At 08:20 Mitesh said it wasn't silly at all, and that it was very close to a comparison Positron prepares internally. He called it "a great advertisement for us," and at 27:40 asked whether it was public so Positron could showcase it in its own internal demos. He also asked where the data came from, which the episode never answered, so treat the on-screen figures (about $10,000 per million on GPU versus $3,600 on Positron) as illustrative, not a benchmark.
Mitesh said energy use was the metric he'd point to first: for the same power, you can push two to five times more tokens, and Positron's systems can go into air-cooled data centers where new GPUs can't.
Don't present it as GPUs losing everywhere. Mitesh was careful here: Nvidia's architecture works for training and most inference. Positron targets specific cases (very large models, long context, image and video generation with large outputs, and code generation, where he said about 90% of tokens are cached). That's where he claims two to four times better performance. A next version would let you pick those workloads explicitly and show where GPUs still win.
For more on the hardware side, see our AI infrastructure and semiconductors pages.
Sam's simulator opens with the claim that inference, not training, is where about 90% of real-world AI costs happen, and that GPUs were built for graphics rather than text generation. Mitesh Agrawal of Positron added that energy and chip supply are the real-world limits.
Sam's version is an explainer plus a simulation: pick a workload (chat prompts or image generation), raise the scale, watch effective margin and cost per million tokens on GPU infrastructure, then switch the same workload to Positron. The episode didn't cover what tools he built it with.
No. Mitesh said Positron's advantage is for specific workloads: very large models, long context, and image or video generation where outputs need a lot of memory. Nvidia's architecture still works for most inference.
The episode didn't say. Mitesh said the comparison was very close to Positron's own internal one but asked where the data came from, so treat the on-screen figures as illustrative.