Built This Week/AI ToolsAI Tools
gpt-oss
gpt-oss-120b felt close to GPT-4.1 in our one live test and ran in seconds on GroqCloud; its real promise is near-zero cost for always-on agents.
By Jordan Metzner & Sam Nadler · From our episodes · Updated October 2026
Our verdict on gpt-oss
gpt-oss is OpenAI's open-weight model family, and the 120b version impressed Jordan the week it came out. On GroqCloud it wrote ten promo emails in about five seconds, and he put its quality near GPT-4.1, "or even close to o three."
The honest caveat: we only ran one live test in August 2025 and have not built a product on it since. Treat this as a first look, not a long-term review.
What we did with gpt-oss
We have no build page for gpt-oss. What we have is the demo in Episode 7, recorded the day after OpenAI released gpt-oss-120b and gpt-oss-20b:
- The setup. Jordan opened GroqCloud, picked the 120b model and set a system prompt: "you're a marketing expert."
- The prompt. Write ten emails announcing advertising for Built This Week, a podcast the model probably had never heard of.
- The result. Ten emails in 5.3 seconds. Jordan added "no dashes" and it rewrote all ten in 4.2 seconds. "Maybe they're not perfect," he said, but the speed was the point.
And what this means is these are models that have essentially zero cost to run.
— Jordan Metzner, Episode 7
How to use gpt-oss
These steps come from Jordan's demo and from how he described using the models at 16:27 and 17:05:
- Pick the size for your hardware. Jordan said the 20b model runs on most computers. The 120b model is the stronger one, but you will want a host.
- Use a fast host for the big model. He ran 120b on GroqCloud and mentioned AWS had announced support too. You pay for compute, not for OpenAI's API.
- Set a role in the system prompt. "You're a marketing expert" was enough for the email test.
- Iterate with tiny edits. Because a full rerun takes seconds, small fixes like "no dashes" cost almost nothing.
- Run it all the time, not on demand. Jordan's bigger idea: one server reading every line of your code base all night looking for bugs, another working on documentation, and starting over when it finishes.
Why OpenAI gave it away
Sam asked Jordan what the release meant strategically. His answer at 19:48: it builds a moat. If your commercial model is not better than OpenAI's free one, "why would anyone talk to you?" Giving these away raises the floor for everyone and makes OpenAI's paid models look like the premium tier.
Where gpt-oss falls short
- It is not OpenAI's best model. Jordan's read was that this is deliberate. Anyone selling a commercial model now has to beat OpenAI's free one, while OpenAI keeps its premium models a step ahead.
- Output still needs editing. The emails were drafts, not finished copy, and they used dashes until told not to.
- We have not stress-tested it. We did not try it on code, long documents or agents, and we did not measure hosting cost. Our spend across models was low enough in April 2026 that Jordan said we just use "whatever the best model is today."
- Speed costs money. Positron's Mitesh Agrawal used gpt-oss as his example in Episode 37: if you want thousands of tokens per second per user, you pay more because it takes more chips. See AI infrastructure for more on that trade-off.
And, eventually, these open source models, even now, are are good enough for many tasks.
— Jordan Metzner, Episode 7
gpt-oss compared
We have no full comparison page yet. The one comparison on air: Jordan said gpt-oss-120b is roughly equivalent to GPT-4.1 but runs on much cheaper infrastructure. For the closed models we use day to day, see ChatGPT and ChatGPT vs Claude.
Episodes featuring gpt-oss
- Episode 7: release news, the GroqCloud speed demo, and what it means for OpenAI's strategy.
- Episode 13: Jordan mentions running OpenAI's open models on GroqCloud.
- Episode 37: gpt-oss as the example for paying more for tokens per second.
FAQ
What is gpt-oss?
OpenAI's open-weight models, released in August 2025 as gpt-oss-120b and gpt-oss-20b. Unlike OpenAI's earlier models, you can host them yourself and only pay for the servers.
Is gpt-oss any good?
Jordan called the 120b model roughly equivalent to GPT-4.1, even close to o3, in Episode 7. Our test was short: ten marketing emails generated in 5.3 seconds on GroqCloud, then regenerated in 4.2 seconds.
Can gpt-oss run on my own computer?
Jordan said the 20b model can run on most computers, so you could have it work all night on your own machine. The 120b model needs more powerful but still cheap infrastructure such as GroqCloud.
What is gpt-oss good for?
Jobs where cost matters more than peak quality: an agent that reads your whole code base all night looking for bugs, another that writes documentation, or fast bulk drafts like marketing emails.