Built This Week/AI ToolsAI Tools
Claude Code
Claude Code is the fastest way we know to go from a PRD to a working app, especially with several sessions running in parallel, but we hand finished work to Codex to catch the bugs it misses.
By Jordan Metzner & Sam Nadler · From our episodes · Updated October 2026
Our verdict on Claude Code
Claude Code is the tool we reach for when we want to go from a written spec to a working app fast. Sam, who has never written code, built most of his internal tools with it. Its weak spot is that it can be confidently wrong, so for polish and bug hunting we now often switch to OpenAI Codex.
What we built with Claude Code
- Sam's ATS (August 2025): Sam's first time using Claude Code. He built a full applicant tracking system front end, with jobs, candidates, a Kanban pipeline, offers and analytics, in about an hour. Some buttons did not work yet.
- Kalshi prediction market trading bot (January 2026): Jordan built a full-stack bot with Opus 4.5, Cursor and Claude Code. It won 8 of 12 trades and still lost about $7 to fees.
- ScreenEval (January 2026): a recruiter screen evaluator and coaching tool. Sam built it in about six hours of work between meetings, in the Claude Code terminal with Supabase.
- Personal DNA and health insights dashboard (February 2026): Jordan gave Claude Code his raw 23andMe file, blood work and medications. It runs locally on SQLite and even wrote a report for his sister's doctor.
- Candidate highlight reel builder and Offsiteio estimator (March and April 2026): both started in Claude Code, then moved to Codex for bug fixes. The estimator cut quote turnaround from about a week and a half to three or four days.
- Internal web app redesign (April 2026): Sam designed it in Claude Design, handed the zip to Claude Code, and had the new design live in production in about thirty minutes.
How we use Claude Code
- Write a PRD first. Sam's ATS started from a generic PRD written in ChatGPT. A second PRD described exactly how our recruiters work and which features of our current ATS they liked (Episode 9, at 03:15).
- Split the work across parallel sessions. Give each session one page or feature. Sam's practical limit was three, and he had to stay engaged with each one.
- Prefer the terminal once you are comfortable. You can open as many Claude Code sessions as you want there. Claude Cowork is a good starting point for non-technical builders, because you point it at a folder of docs and it shows the plan and progress.
- Lock the front end before the backend. Jordan's rule from Episode 9: iterate the screens with real users first, then build the database.
- Keep the stack simple. ScreenEval and the Offsiteio estimator use Supabase for data and deploy through AWS Amplify and GitHub.
- Keep sensitive data local. Jordan's DNA dashboard never leaves his machine.
- Reuse skills. Jordan keeps a skill with his preferred design features and layouts and shares it with Sam (Episode 29, at 14:12).
- Hand off for polish. When the app works, run Codex on high reasoning to find the bugs Claude Code was sure were not there.
I think, you know, for me, personally, my limit was around three agents, making sure I was, like, you know, engaging appropriately with each agent on what it was building.
— Sam Nadler, Episode 9
Where Claude Code falls short
- Confident bugs. Jordan's main reason for moving his own work to Codex in March 2026.
- Cost. Anand Chandrasekaran of Arya Health said in Episode 24 that it is "already very expensive." Ben Lerner of Espresso AI said the opposite in value terms: it costs far less than a third of an engineer.
- Serious backend work. Ben said 25 Claude Code agents running backend infrastructure in the background is "just not there yet" (Episode 27).
- Code review. Bito's Amar Goel said asking Claude Code to review your code catches much shallower issues than a dedicated reviewer (Episode 23).
- Non-technical users. Adir Ben-Yehuda of Autonomy AI compared CLI agents like Claude Code to MS-DOS for product managers (Episode 43).
- Mission-critical code. Code Metal's Ryan Aytay said AI-written code is not safe enough on its own for systems at the edge (Episode 48).
One thing I would notice like repetitively with with Claude Code is that it would make mistakes, but be confident that there were no mistakes.
— Jordan Metzner, Episode 35
Claude Code compared
Episodes featuring Claude Code
- Episode 9: Sam's first Claude Code build, with three parallel agents.
- Episode 26: the Kalshi bot, plus Jordan on plugins and observability tools built on top of Claude Code.
- Episode 29: ScreenEval, and Claude Code in the terminal vs Cowork.
- Episode 30: the DNA dashboard and Agent Teams.
- Episode 35: Claude Code vs Codex vs Cursor with Juma's Iliya Valchanov.
- Episode 40: the Claude Design handoff to Claude Code.
FAQ
How do you use Claude Code?
We start with a PRD, give it to Claude Code, then run a few sessions in parallel, each on its own page or feature. Sam built a full ATS front end this way in about an hour on his first try.
Can a non-developer use Claude Code?
Yes. Sam has never written a line of code and built ScreenEval, a full-stack recruiter coaching app, in about six hours with Claude Code and Supabase. Product people who do not know the command line still find it hard, according to Autonomy AI's Adir Ben-Yehuda.
Is Claude Code worth it?
For us, yes. Espresso AI's Ben Lerner said his team pays far less for Claude Code than a third of an engineer, and estimated everyone is 10 to 30% more productive. One guest, Anand Chandrasekaran, called it very expensive compared with everything else.
Should I use Claude Code in the terminal or Claude Cowork?
Cowork is a friendly entry point with a polished plan and progress view. The terminal is faster because you can run as many Claude Code sessions at once as you want.
What is Claude Code bad at?
Jordan's main complaint is that it sometimes makes mistakes while being confident there are none. Guests also said it is not ready for serious backend infrastructure run by dozens of agents, or for mission-critical code on its own.