**Julian Goldie** (0:00)
Today we have Claude Fable 5 and new announcements on Mythos. So two new model announcements today. This is directly from Claude. So Claude Fable 5, which is a Mythos class model, they've basically made safe for general use. And these are two of the most powerful models ever released from Claude. So I'm going to show you exactly what they mean, what's happened, what the differences are, how it works, etc. We've already plugged it into our agent operating system. As you can see here, it says I'm powered by Claude Fable 5, which is the most recent available model from Claude. And so basically what they've said is today we are launching Claude Fable 5, which is a Mythos class model. So basically they've taken Mythos and the idea of Mythos, which is the model they couldn't release to the public, and they kept closed to Project Glasswing only. And then what they've done from there is they've created Fable 5, which is still in the same class. And so this is better than any model ever created, basically state of the art on all the benchmarks previously. So it's great at software engineering, knowledge work, vision, scientific research, many other areas. The longer and the more complex the task, the larger Fable 5 lead is over these models. Here's an example of what you can build with this. Here's an example guide that we actually created about Claude Fable 5 and Mythos 5
And this was generated fully with the setup, right, with Claude Fable 5 So as you went out, it did all the research. It took quotes, as you can see here, organized some images as well on the page using Chat GPT image API. And then you can see it's actually organized all of the research together. So what actually dropped? Why is this useful, et cetera? Let's take a look at this. So here's some of the biggest announcements. Number one, it's a brand new model class. It's not just like a new version. So for example, for years, you really had Claude in three sizes, which is Haiku, Sonnet, Opus, Sonnet and Opus 4.8 just came out last month. Now there's a fourth class sitting above all, and that is Mythos. Now Claude Fable 5 is the first Mythos class model the public has ever been allowed to touch. So Fable 5 and Mythos 5 are the same underlying model in many ways. The only difference is safeguards. So Fable 5 hasn't. Mythos 5 isn't safe, and that's why it's only restricted to vetted partners. Having said that, they have said that it will be released in the future.
So we'll keep talking about that in a second. These are the benchmarks here. So if we compare, for example, Opus 4.8 versus Fable 5 on WE Bench Verified, you can see that it scores 95 percent, this 88.6 percent. It's a big jump in benchmarks right here. You might be saying, okay, you don't care about the benchmarks. I agree, it's always better to test this stuff yourself, but it is interesting to pay attention to this and to understand what the differences are. You can also see SW Bench Pro here. So Fable 5 scored 80 percent, Opus 4.8 scored 69 percent, and GPT 5.5 scored 58.6 percent. So Fable 5 is a huge step up in SW Bench, which is software engineering as a benchmark. You can see for Frontier Code, the Fable 5 scores 29.3 versus Opus 4.8, which is 13.4. So it's pretty much doubled or more than doubled on that benchmark, which is pretty wild. It's also the first model ever to break 90 percent on Hexi's Core Analytics benchmark. So this is a 10 point jump over. So pretty wild stuff right there. It's built for long horizon real work. So it has a million token context window, 28k max output per request. On long context retrieval, the Mythos class model scored 79.4 versus Opus' 68.1. And tool Athlon, which is a tool use marathon. So these are agented tasks that are basically looking at, OK, how powerful is it? Can it basically run directly with something on a long horizon task? Fable 5 scored higher than Opus 4.8 and finished in fewer turns. So it scored higher and finished in fewer turns. So 61.7% in 19.8 average turns versus 59.9% in 24.5 turns for Opus 4.8. So number one, it scored better. And number two, it scored better in less turns, which is pretty amazing when you think about it. So you get better results, fewer steps. I will say, in theory, that means cheaper agent runs, but actually, in reality, the cost of the API is much higher for Fable 5 If you haven't seen the costs already, let me show you this. So these are, where are they? We'll come on to the costs in a minute. And now this is pretty interesting. So if you're wondering, okay, what sort of crazy stuff can I do? It can play Pokemon FireRed using vision alone, reading the screen. Previous models needed special harnesses. Anthropics internal experts used it for some interest in research and design, improving their results by roughly, it ran some novel research autonomously for a week. And in actual blind tests, biologists preferred its novel molecular hypotheses 80% of the time over Opus class models. Stripe have also used it as well on a big 50 million line codebase migration too.
9 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000772166311