Built This Week/Guests

Guests

Carina

Carina founded Axiom Math, which is building a self-improving AI mathematician that conjectures, proves and auto-formalizes math into Lean so every proof can be checked by running it.

Carina is the founder and CEO of Axiom Math, which is building a self-improving AI mathematician. Its system proves theorems, proposes new problems and auto-formalizes math into the Lean programming language, so every proof can be checked by a computer. She joined us in Episode 22 (November 21, 2025), the week Gemini 3 came out.

What Axiom Math does

Carina's goal at 09:41 is a model that does more than score well on Olympiad benchmarks. She wants it to work like a research mathematician: conjecturing new, unseen problems, then trying to prove and verify them. A prover and a conjecturer talk to each other in a self-play loop.

The same system auto-formalizes, turning the huge corpus of written math into Lean. Her bet is that reinforcement learning, which has driven big jumps in verifiable domains like coding, can do the same for serious mathematics once math is code.

Key ideas from the conversation

  • Data, not compute, is the bottleneck (11:14). Axiom uses a lot of both GPUs and CPUs: language models run on GPUs and Lean compiles on CPUs. The hard part is that there just isn't much Lean.
  • Synthetic data over armies of labelers (12:24). Some companies hire large teams of Lean programmers to write data by hand, which is expensive and doesn't scale. Axiom converts informal math it scrapes into Lean and finds ways to grow its seed of formal math many times over.
  • Mathlib started in a classroom (13:27). Kevin Buzzard at Imperial College London had freshmen formalize proofs in Lean. Bored advanced students ended up formalizing the undergraduate math corpus, and two or three of them now work with Axiom. Carina met several through an MIT and Imperial College London exchange around 2019.
  • Verified, not judged (15:40). An LLM proof in English can hide a bug in 3,000 lines. Axiom's output is an executable Lean program.
  • On Gemini 3 (18:34). Her team likes it and finds it much cheaper to run than GPT. She also said Axiom's model beat Gemini 3 on a math proving benchmark she didn't name.
And there are more than 1,000,000,000,000 tokens of Python code, only about, like, 10,000,000 tokens of lean code. That's a 100,000 times data gap.

— Carina, Episode 22

We just run it like you run a Python computer program. I think that's the biggest difference between these formal models and the large language models.

— Carina, Episode 22

What we built for the episode

Jordan built the LLM Math Roaster with Axiom in mind. It sends a math problem to Gemini 2.5 Pro, GPT-5, Claude Sonnet 4.5 and Grok 4 Fast, asks each for a Lean 4 proof and has ChatGPT judge the results on a leaderboard. It also has run history, custom problems and an API so Axiom's team can submit problems.

Every model passed a 2 + 2 = 4 sanity check. Carina then picked Fermat's Little Theorem. The judge gave Gemini 98 and ChatGPT 70, and Carina spotted that one model had proved it for natural numbers instead of integers. Her suggestion was to skip the LLM judge entirely.

We could also, right, like take that lean proof and then put it in like lean land and like, you know, and see if it compiles, which should be hopefully be pain more painless than having LM judge.

— Carina, Episode 22

When Jordan said running thousands of problems would expose each model's weak spots, Carina said that was the first thing Axiom did, back when the office was camper chairs and a plastic folding table.

In the news segment

We covered Gemini 3 and Google Antigravity, Jeff Bezos' Project Prometheus and Suno's $250M raise. Carina was excited that Prometheus looks focused on AI for physical science. On Suno, she imagined people with mathematical creativity but no proof training one day using Axiom's tools to produce real theorems, the way non-musicians make songs. More on AI and learning: AI in education.

FAQ

What is Axiom Math building?

A self-improving AI mathematician. A prover and a conjecturer talk to each other in a self-play loop, and the same system auto-formalizes existing math into the Lean programming language so proofs can be verified programmatically.

Who is the founder of Axiom Math?

Carina is the founder and CEO of Axiom Math. She joined Built This Week in Episode 22, in November 2025, to talk about AI for math and Lean.

Why is Lean training data scarce?

Carina says there are more than 1 trillion tokens of Python code but only about 10 million tokens of Lean, a 100,000x gap. Axiom generates synthetic data, including auto-formalizing informal math into Lean.

How is Axiom Math different from ChatGPT at math?

LLMs write proofs in natural language, which is hard to check step by step. Axiom's hybrid prover outputs Lean programs that are executed and verified, so no LLM judge or human judge is needed.

Where is Lean taught?

Carina named the University of Washington, CMU, Stanford's automated reasoning lab and several UK universities. Mathlib started as a teaching effort by Kevin Buzzard at Imperial College London.