Built This Week/AI Tools

AI Tools

Grok

We have tested Grok as one model among several, and guests respect how fast xAI caught up, but as of 2026 we have not found a high-quality use case to put it in production.

Our verdict on Grok

Grok is a model we test, not one we run. In February 2026 Jordan said plainly that we use Anthropic, OpenAI and Gemini models in production, but not Grok, because we haven't found a high-quality use case for it. Our hands-on evidence is thin: two builds where Grok was one model among several.

Guests are kinder. Amar Goel of Bito called it miraculous how fast xAI built a decent model with an impressive cost-performance ratio.

Like, it's miraculous to me how quickly they built, like, a pretty decent model, you know, with Grok four and Grok fast code one.

— Amar Goel, Episode 23

What we built with Grok

  • LLM Math Roaster (November 2025): Jordan built it with Carina of Axiom Math in mind. It sends a math problem to Gemini 2.5 Pro, GPT-5, Claude Sonnet 4.5 and Grok 4 Fast, asks each for a Lean 4 proof, and has ChatGPT judge them. On Fermat's Little Theorem the judge gave Gemini the best score; the episode doesn't call out Grok's result.
  • Solopreneur finance dashboard for Collective (September 2025): Jordan started the cash-flow forecaster on a roughly 120B model "just because I had access to it". The episode notes call it Grok 120B; the audio is hard to make out. Either way it couldn't handle the volume of P&L and balance sheet data, so he moved to Gemini.

How we use Grok

  1. Put it in a lineup, not on its own. The Math Roaster runs the same problem through four models side by side, with response times, so you see where Grok stands on your task.
  2. Judge with something other than an LLM if you can. Carina pointed out that compiling the Lean proofs is more reliable than an LLM judge. The same goes for any model comparison.
  3. Test with your real data volume early. The Collective forecaster failed on the size of real financial statements, not on a toy example. Check context limits before you build around a model.
  4. Check API access before you plan. Jordan wanted the newest Grok in the Math Roaster but didn't have API access in time.
  5. Offer it as a user choice. Emblem's August Kiles said they only expose xAI in a dropdown for users who already prefer it; their backend doesn't use it.
  6. Keep production on what has proven itself. For us that is Claude, OpenAI and Gemini, with Claude Code and OpenAI Codex for coding.

Where Grok falls short

  • Coding. Jordan's August 2025 take is below. In April 2026 he added that Elon has said he is aware Grok is behind in enterprise coding.
  • Safety. When Grok 4 launched in July 2025, Jordan found its performance impressive but thought xAI probably cut some corners on safety, and joked that the first thing Elon launched was an NSFW companion mode.
  • Privacy. In August 2025 public Grok share links got indexed by Google, a week after the same thing happened to ChatGPT.
  • Adoption. In October 2026 Jordan said xAI's newest model hadn't caught on with Twitter or Reddit yet, while adding that you shouldn't count out Elon and the Cursor team.
It's okay. It's not better than anything else. I mean, it's not better than than Claude for coding, or, you know, GPT five, I think, now.

— Jordan Metzner, Episode 9

We haven't really found them of, like, high quality of use case yet.

— Jordan Metzner, Episode 31

Grok compared

  • Gemini vs Grok: Gemini took in the large financial files the 120B model couldn't, and won the Math Roaster's Fermat test.
  • Claude vs Grok: for coding, Jordan rates Claude higher.
  • In April 2026 Jordan ranked Grok third or fourth among AI labs, and said xAI's compute deal with Cursor may put it back in the number three slot against Google.

Episodes featuring Grok

FAQ

Is Grok good for coding?

Not in our experience. In August 2025 Jordan said it was okay but not better than Claude for coding or GPT-5, and in April 2026 he noted Elon himself has said Grok is behind in enterprise coding.

Grok vs Gemini: which should I use?

For our Collective finance dashboard Jordan started with a roughly 120B model the episode notes call Grok, but it could not take in the P&L and balance sheet data, so he switched to Gemini. In the LLM Math Roaster the judge gave Gemini the top score.

Do companies use Grok in production?

We don't. Emblem's August Kiles said they only offer xAI as a dropdown option for users who prefer it and don't use it in their backend.

Is Grok worth trying?

As one model in a lineup, yes. Bito CEO Amar Goel called it miraculous how quickly xAI built a decent model with a strong cost-performance ratio, and Jordan said in October 2026 not to count out Elon and the Cursor team.