Carina is the founder and CEO of Axiom Math, which is building a self-improving AI mathematician. Its system proves theorems, proposes new problems and auto-formalizes math into the Lean programming language, so every proof can be checked by a computer. She joined us in Episode 22 (November 21, 2025), the week Gemini 3 came out.
Carina's goal at 09:41 is a model that does more than score well on Olympiad benchmarks. She wants it to work like a research mathematician: conjecturing new, unseen problems, then trying to prove and verify them. A prover and a conjecturer talk to each other in a self-play loop.
The same system auto-formalizes, turning the huge corpus of written math into Lean. Her bet is that reinforcement learning, which has driven big jumps in verifiable domains like coding, can do the same for serious mathematics once math is code.
And there are more than 1,000,000,000,000 tokens of Python code, only about, like, 10,000,000 tokens of lean code. That's a 100,000 times data gap.
— Carina, Episode 22
We just run it like you run a Python computer program. I think that's the biggest difference between these formal models and the large language models.
— Carina, Episode 22
Jordan built the LLM Math Roaster with Axiom in mind. It sends a math problem to Gemini 2.5 Pro, GPT-5, Claude Sonnet 4.5 and Grok 4 Fast, asks each for a Lean 4 proof and has ChatGPT judge the results on a leaderboard. It also has run history, custom problems and an API so Axiom's team can submit problems.
Every model passed a 2 + 2 = 4 sanity check. Carina then picked Fermat's Little Theorem. The judge gave Gemini 98 and ChatGPT 70, and Carina spotted that one model had proved it for natural numbers instead of integers. Her suggestion was to skip the LLM judge entirely.
We could also, right, like take that lean proof and then put it in like lean land and like, you know, and see if it compiles, which should be hopefully be pain more painless than having LM judge.
— Carina, Episode 22
When Jordan said running thousands of problems would expose each model's weak spots, Carina said that was the first thing Axiom did, back when the office was camper chairs and a plastic folding table.
We covered Gemini 3 and Google Antigravity, Jeff Bezos' Project Prometheus and Suno's $250M raise. Carina was excited that Prometheus looks focused on AI for physical science. On Suno, she imagined people with mathematical creativity but no proof training one day using Axiom's tools to produce real theorems, the way non-musicians make songs. More on AI and learning: AI in education.
A self-improving AI mathematician. A prover and a conjecturer talk to each other in a self-play loop, and the same system auto-formalizes existing math into the Lean programming language so proofs can be verified programmatically.
Carina is the founder and CEO of Axiom Math. She joined Built This Week in Episode 22, in November 2025, to talk about AI for math and Lean.
Carina says there are more than 1 trillion tokens of Python code but only about 10 million tokens of Lean, a 100,000x gap. Axiom generates synthetic data, including auto-formalizing informal math into Lean.
LLMs write proofs in natural language, which is hard to check step by step. Axiom's hybrid prover outputs Lean programs that are executed and verified, so no LLM judge or human judge is needed.
Carina named the University of Washington, CMU, Stanford's automated reasoning lab and several UK universities. Mathlib started as a teaching effort by Kevin Buzzard at Imperial College London.