Built This Week/AI ToolsAI Tools
OpenAI Codex
OpenAI Codex on extra high reasoning is slower and less fun than Claude Code, but it writes cleaner code and finds the bugs Claude Code misses, so it is Jordan's daily driver and our polish step.
By Jordan Metzner & Sam Nadler · From our episodes · Updated October 2026
Our verdict on OpenAI Codex
OpenAI Codex is the coding agent we trust when code quality matters more than speed. By March 2026 Jordan had moved his own work to Codex, and by June his daily driver was Codex 5.5 on extra high. Sam moves apps he started in Claude Code into Codex to find bugs and polish them. It is slower and less fun, and its token spend has surprised us.
Our stance has changed more than once. Jordan tried Codex in mid-2025 and called it "a pretty poor experience." A few months later his team had moved to it. We keep saying the same thing: we switch to whatever works best that month.
What we built with OpenAI Codex
- Strudel AI DJ mixer (September 2025): a two-deck browser mixer built on Strudel, with an AI button that writes a new beat for each channel. Jordan built most of it with Codex over Labor Day weekend. It worked, with some bugs.
- Solopreneur finance dashboard (September 2025): first version in Bolt, then moved into Cursor, where Codex coded the front end and wired up the APIs. Gemini handled the large P&L inputs.
- Tix LAX (October 2025): a Mapbox map of Los Angeles parking tickets with an officer leaderboard. Jordan coded "almost all of it" with Cursor and Codex 5.
- Rep-by-rep AI personal trainer (February 2026): Sam brought an existing app into the Codex desktop Mac app and ran GPT-5.2 Codex on extra-high reasoning to find bugs and fill gaps. Still in beta at recording.
- Candidate highlight reel builder (March 2026): started in Claude Code, then "mainly built in codex" with Remotion as the video engine. Recruiters now render a reel in 5 to 10 minutes instead of hours or days.
- Offsiteio estimator (April 2026): started in Claude Code; after some bugs, Sam moved it to Codex for fixes and updates. It runs on Supabase and is used daily.
It helped me build, like, the majority of what you saw inside my DJ tool. It's incredibly well developed.
— Jordan Metzner, Episode 11
How we use OpenAI Codex
- Get a first version up fast. Jordan's Collective dashboard started in Bolt; Sam's reel builder and estimator started in Claude Code. Codex comes in once there is something real to improve.
- Run it where you already work. In 2025 Jordan used Codex inside Cursor, his IDE. In 2026 Sam used the Codex desktop Mac app, which he found very similar to Claude Cowork (Episode 30, at 06:47).
- Use extra high reasoning for bug hunting. Sam's runs took five to ten minutes each, but they found gaps he was not even looking for.
- Drop the reasoning level for quick iterations. Sam's advice at 10:44: a lighter model when you want speed, the deeper models when you are polishing toward production.
- Style by describing, not editing. Sam has zero video editing skills; he shaped the reel builder's template by telling Codex in plain language how the colors and graphics should look.
- Let developers pick. Jordan gives our developers access to all the tools. Some still prefer Claude Code; he and others found Codex much better.
- Watch the bill. Heavy features such as computer use burn tokens fast. In June 2026 we were still debating whether to cap usage or just monitor it and rein it in if it gets out of control (Episode 45, at 21:13).
I took a previously existing product and then brought it over to Codex to run this extra high reasoning 5.2 GPT model, and I've found it great to kind of polish off some of those things
— Sam Nadler, Episode 30
Where OpenAI Codex falls short
- Speed and fun. Jordan says it is "a little bit slower" and "not as fun" as Claude Code, which keeps you entertained while the agent works (Episode 35, at 17:52).
- Token spend. In June 2026 the computer use feature "killed us on spend," and it took about two and a half weeks to find the cause. Sam said charges were coded differently across the console and the Codex console.
- A moving dashboard. Jordan said the Codex dashboard changed almost every day for two weeks.
- Serious backend work. Ben Lerner of Espresso AI uses Codex prompts for dashboards and throwaway pandas scripts, but said models are not good enough yet for serious backend infrastructure (Episode 27, at 13:55).
- Not everyone switched. Iliya Valchanov of Juma said Claude Code works so well his team sees no reason to move.
But the output of the code, both with 5.3 and now especially with 5.4 on extra high, I found to be significantly better quality.
— Jordan Metzner, Episode 35
OpenAI Codex compared
- Claude Code vs OpenAI Codex: Claude Code is faster and more fun; Codex on extra high writes cleaner code and catches more bugs. We use both.
Two guest views add context. Emanuele Melis liked that Codex's browser view clicks through the UI until web work is done (Episode 46). David Petrou said OpenAI has been more welcoming than Anthropic to developers who want to use their subscription outside the official coding app (Episode 51).
Episodes featuring OpenAI Codex
- Episode 11: the Strudel DJ mixer, and why our engineers moved to Codex.
- Episode 14: Codex in Cursor for the Collective dashboard; "we got no loyalty."
- Episode 15: Tix LAX, and Jordan calling Codex 5 the top coding tool.
- Episode 30: Sam reviews the Codex desktop app on extra-high reasoning.
- Episode 35: Jordan on Codex vs Claude Code code quality.
- Episode 36 and Episode 38: two of Sam's tools moved from Claude Code to Codex.
- Episode 45: the computer use spend spike.
- Episode 46: Codex 5.5 extra high as Jordan's daily driver.
FAQ
Is OpenAI Codex good?
For us, yes. Jordan called it underdeveloped in mid-2025, then built most of his Strudel DJ mixer with it over Labor Day weekend. By June 2026 his daily driver was Codex 5.5 on extra high.
Is OpenAI Codex better than Claude Code?
It depends on the job. Jordan says Claude Code is faster and more fun but sometimes confidently wrong, while Codex 5.3 and 5.4 on extra high gave him significantly better code quality. Sam uses both and moves apps to Codex to polish them.
How long does OpenAI Codex take to run?
On the extra-high reasoning setting, Sam's runs usually took five to eight minutes, sometimes ten. That is why we use faster settings for quick iteration and save extra high for polishing.
How much does OpenAI Codex cost?
We did not track a price on the show. In September 2025 Jordan said OpenAI was subsidizing it and his team was maxing out usage; in June 2026 the computer use feature drove a token spend spike that took about two and a half weeks to trace.
What is OpenAI Codex bad at?
It is slow on deep reasoning settings, the billing dashboards were hard to read, and Espresso AI's Ben Lerner said models like it are not yet good enough for serious backend infrastructure.