Built This Week/AI by Industry

AI by Industry

AI in Semiconductors

Our two semiconductor guests work both directions: Positron designs chips for AI inference, and Vinci uses an AI physics model to help engineers design chips and packages.

Built This Week has covered AI in semiconductors in two episodes, from opposite directions. Mitesh Agrawal, CEO of Positron AI, designs silicon for AI inference. Hardik Kabaria, co-founder and CEO of Vinci, uses an AI physics model to help engineers design chips, packages and boards. Both said the hard limits are physical: memory supply, heat, warpage and how long each design loop takes.

Now instead that engineer, she might be performing 10 to 20 analysis per day, exploring different configurations that came to her from design. With us, she's able to do thousands.

— Hardik Kabaria, Episode 50

How companies we talked to use AI in semiconductors

Vinci: a physics model for chip and package design

Vinci's demo ran warpage analysis on a native semiconductor package design at three-micron resolution and a 225°C bonding temperature: 1.2 billion degrees of freedom in under four minutes. Warpage decides whether a package ships, because bonds only form within micron tolerances, and Hardik's demo noted that some advanced products had been redesigned because warpage wasn't predicted early enough. The model behind it has half a billion parameters, trained from scratch for thermal and thermo-mechanical physics. (Episode 50)

Hardik said the bigger change is who gets answers. A designer or power engineer can ask whether a CPU can still be air-cooled without waiting for a separate simulation team, which can cut loops of weeks. Vinci deploys behind customers' firewalls and works with memory, fabless, foundry and equipment companies, mostly tier-one enterprises it can't name.

Positron AI: inference chips built on commodity memory

Positron designs inference silicon around memory bandwidth utilization and capacity, with up to two terabytes per card. Mitesh said its first-generation product proved the bandwidth part, and the second chip tapes out later in 2026, arriving in mid-2027. Its biggest design choice is avoiding HBM, the memory Nvidia uses, and building on the commodity memory found in phones and laptops. (Episode 37)

Can you fabricate it? You know, can you actually get the components to it? And if if I'm using the same memory components as NVIDIA, like, good luck.

— Mitesh Agrawal, Episode 37

Mitesh was careful about Nvidia, calling it one of the world's smartest companies. Positron goes after a few workloads, such as very large models, long contexts and image or video generation, where he says it can beat the latest Nvidia chips by two to four times. He also said Positron was one of the first accelerators integrated with Arm's first CPU of its own, and that it has reached tens of millions of dollars in revenue. More on the cost side of running models is on our AI infrastructure page.

What we built

  • AI inference cost simulator (GPU vs Positron): Sam made this explainer for the Positron episode. It walks through why GPUs, designed for graphics, waste cycles on text generation, then compares a GPU fleet with Positron hardware on latency, energy use and cost per million tokens as the workload grows. Mitesh called it close to the comparisons Positron prepares in-house, but said his chips only win on memory-heavy jobs.

We didn't build anything for the Vinci episode. Jordan compared Vinci to AI design tools: just as those let engineers who aren't designers build front ends, Vinci lets people who aren't simulation engineers run physics checks.

What's still hard

  • Memory supply. Mitesh said HBM is the best memory but a startup can't get it. Even commodity memory contract prices rose two to two and a half times, though allocation and capacity are there.
  • Slow, scattered design loops. Design, performance checks and manufacturability often sit with different teams, sometimes different companies, so iterations take weeks. A tape-out can cost millions of dollars and months.
  • Long jobs and too much data. Some Vinci customers ran queries that took seven days, so the system has to survive compute failures. The output is gigabytes, and Hardik said people decide on simple charts, so the product has to summarize.
  • More than one kind of physics. Customers want electromagnetics and RF next, and Hardik said no workflow today ties all the disciplines together.
  • Design provenance. Hardik expects guardrails and ways to tell who generated a design, because AI-generated designs may end up driving real machines.

Episodes on AI in semiconductors

FAQ

How is AI used in semiconductor design?

Vinci, which we hosted in July 2026, trained a half-billion-parameter physics model for thermal and thermo-mechanical analysis. A thermal engineer who ran 10 to 20 analyses a day can run thousands, and chip designers can get physics answers without waiting for a specialist.

What is warpage in semiconductor packaging?

Warpage is how much a package bends during assembly. Bonds only form when it stays within tolerances measured in microns. Vinci's demo ran a manufacturing-resolution warpage analysis with 1.2 billion degrees of freedom in under four minutes.

Why do AI inference chips need so much memory?

Positron CEO Mitesh Agrawal says very large models, long contexts and image or video generation need far more memory than a GPU's few hundred gigabytes. Positron builds cards with up to two terabytes using commodity memory instead of HBM.

Can startups compete with Nvidia on AI chips?

Mitesh calls Nvidia one of the world's smartest companies and says it wins for training and most inference. Positron targets specific inference workloads where he says it can beat the latest Nvidia chips by two to four times.