For Episode 22 (November 21, 2025), Jordan vibe coded the LLM Math Roaster with our guest's company, Axiom Math, in mind. It sends a math problem to four models, asks each for a formal proof in Lean 4, has ChatGPT judge the results, and ranks them on a leaderboard. He called it "productionalized" and gave it history, custom problems and an API. He didn't say how long it took.
On our first call with Axiom founder and CEO Carina, we talked about vibe-coded apps that could be fun for math and useful to her team. Axiom is building an AI mathematician that writes proofs in Lean, so a tool that pits models against each other on Lean proofs was a natural fit. Jordan was upfront: "I'm okay at math but I'm certainly not your level."
Jordan wanted Grok 4.1, Gemini 3 and GPT-5.1 too, but they had just launched and he didn't have API access in time. The episode didn't name the coding tool he used.
So I chose ChatGPT. So ChatGPT will then evaluate the responses from all of the models.
— Jordan Metzner, Episode 22
All four models passed 2 + 2 = 4. Carina then picked Fermat's Little Theorem: for a prime p and any integer a not divisible by p, a^(p−1) ≡ 1 mod p.
So I have a leaderboard. It gave Gemini the best score. ChatGPT is 70, and so on and so forth.
— Jordan Metzner, Episode 22
Gemini scored 98. Carina spotted a problem at 08:22: one answer proved the theorem for natural numbers instead of integers, skipping the negatives. "Just a close miss," she said. That's one problem with one judge, so don't read it as a ranking of the models.
For more model matchups, see Claude vs Gemini and ChatGPT vs Gemini.
In our one real test, Fermat's Little Theorem in Lean 4, the ChatGPT judge gave Gemini 2.5 Pro the top score (98) and ChatGPT's model a 70. That's a single problem with an LLM judge, not a benchmark.
Send the same problem to each model asking for a Lean 4 proof, show the responses side by side with response times, then score them. Jordan used ChatGPT as the judge; Axiom Math's Carina suggested compiling each proof in Lean instead.
It's a weak basis of truth. Carina pointed out that compiling the Lean proof would be a more reliable check than asking another model, and one high-scoring answer had only proved the theorem for natural numbers, not integers.
Jordan's app has a problem set of about 10 problems, custom problem submission, run history, a leaderboard from the judge's scores, and an API with user-generated keys so a team can submit problems programmatically.