GPT-5.6 Sol vs. Claude Fable: Why OpenAI’s new model crushes my benchmark artwork

GPT-5.6 Sol vs. Claude Fable: Why OpenAI’s new model crushes my benchmark

How I AI

July 9, 2026

GPT-5.6 Sol is back, and I ran it through my full How I AI vibe benchmark against GPT-5.6 Terra, Luna, Claude Fable 5, and Sonnet 5 across five categories: PRDs, prototypes, wireframes, debugging, and agentic voice.
Speakers: Claire Vo
**Claire Vo** (0:00)
I have been very, very, very sad the last week because for the last week, I have not had access to my true favorite, top-of-the-line model, GPT-56. But guess what, babes, it is back and I am here to walk you through GPT-56 Sol, GPT-56 Luna, GPT-56 Terra. I'm going to tell you what are these models, how have I been using them, why are they my hearts favorite, and is Fable better than all of them or not? I have been testing this model for a couple of weeks. There was a few days there where we didn't have access and I found myself desperate to get this workhorse model back. Now, we're not just relying on my own opinion. We are going to run the very famous, very new How I AI Vibe Review benchmark against common tasks from PRD writing to prototyping, to whether or not it's cute in my OpenClaw agent. And I'm going to tell you very scientifically if this is the model that you should be working with all the time now. Let's get to it. Okay, you all can read these blog posts, so I'm not going to go into too much depth about the models and the benchmarks. I'll just give you the hits. First, OpenAI is releasing three new versions of their GPT 5.6 model. Sol, which is the next-generation frontier model, the brainiest of the brainiest. Terra, which is a balanced model for efficient everyday work, and Luna, which is akin to their mini or nano models, which is cheap and affordable for high-volume work. You're going to have these three versions of the models. I don't know if these beautiful images are exactly how we should think about the relative capabilities, this big sun, this medium earth, and this tiny moon. But I will say my love letter that is this podcast today is written directly to GPT-56 Sol. This big model is the one I love. Now I have tested Terra and Luna, so I will give you my input there. But really this is going to be all about Sol versus Fable and which one I would use for the type of work that I'm doing every day. Okay, quick note on pricing. Sol is a lot more affordable than Fable. So it's $5 per million input tokens, $30 per million output tokens. I believe Fable, at the time I'm recording this, is 10 on a million input tokens and 50 on a million output tokens. Now, again, you're going to get a little bit of subscription usage built into your OpenAI subscription. So you are going to get a decent amount that you can test with and use. There's been some challenges with the Fable rollout. They've limited when it's been included in the subscription, and so it was supposed to be available till early this week. I think they extended that a little bit at Anthropic, so subscription Claude users could use Fable under their subscription. So we have to see how much Sol usage we get and if, like Anthropic, they're going to take Sol out of the subscription. I suspect not. I suspect this is a model they want people to use. I also suspect this might put pressure on Anthropic to put Fable back into the Claude subscription. But for now, it's more affordable even at API pricing. Now, I'm not going to read through all the benchmarks for you. You can go to this OpenAI blog and read them for yourselves. All I will say is it is the brand new state-of-the-art model from OpenAI. It is the highest performing when using the Ultra mode on Terminal Bench 2.1.
Then they've also evaled it against a couple of cybersecurity benches. So I do think as we get these smarter models, we're going to see a lot more evals and benchmarks around exploits and security. And then very similar to what we're seeing with Fable, there's a lot of conversation in this blog post about the safeguards and security frameworks around the release of this model. I do believe like Fable, it's going to fail over in some tasks that are maybe a little bit riskier, but I have not run into that myself. Now let's get back to how I eval these models. If you missed my episode on Fable, I got kind of bored of the vibey vibe check and I built a extremely scientific How I AI benchmark. Now this How I AI benchmark tests basically a couple of things. It tests the ability to generate good PRDs. It tests the ability for it to wireframe against a couple of different app ideas, develop fully designed, robust designed prototypes, debug code and then talk to me like a human, which is the thing that I care about the most. I'm just going to remind you how I did these benchmarks and then scan you through a couple of the outputs. Since I know what the models are now after I've done the grading, I can show you which ones map to Fable and GPT-56.

26 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000776152330