**Julian Goldie** (0:00)
So, today we have the release of Claude Sonnet 5, apparently the most agentic model from Claude and Anthropic yet. And apparently, you can see the announcement here, just dropped a few hours ago. It can make plans, use tools like browsers, terminals run autonomously, level. Just a few months ago, required larger and more expensive models. We'll come on to this in a minute. I'm not going to hype it up here because you'll see from my test.
I'm just going to tell you the honest truth. So here you can see you got Sonnet 5, you got Sonnet 4.6. So it is a step up from Sonnet 4.6. If you actually look like Sonnet 5, nowhere near the same level as Opus 4.8. So if we look at agentic coding, 63% versus 69% for Opus 4.8. I mean, literally on every single benchmark here, Opus 4.8 is crushing it. Now, if you were comparing 4.6 from Sonnet versus 5, then of course, it's a big step up. But I think for most people like you, you're probably still going to stick to Opus 4.8 for building stuff, especially at this point. Now, we can expect Fable 5 to drop very soon. Anthropic just announced that they're going to be releasing and giving people access back again within 24 hours for Fable 5 So this kind of just feels like a weird release and I don't know if many people care about it. I mean, if you're watching this, you use Sonnet 4.6 already, are you not just sticking to Opus 4.8? Are you not going to just switch to Fable 5? For me personally, I just don't know if this is really a big update or not. Now, if we actually have a look, what we've created so far on Goldie Bench with Claude Sonnet, let's take a look right here. So if we pull up the benchmarks over here, he did create some pretty cool stuff and I'll show you what we've built, how it works, et cetera, how it performs. So we can gather some examples side by side over here. Here's a couple of examples. So this was like a ray caster maze. I mean, it can create pretty nice stuff, right? Like this is pretty cool.
The graphics, pretty nice, the colors, pretty nice. I mean, it works perfectly as a game. For some reason, the controls are a bit backwards when I test out personally. But apart from that, we're all good here and it's working nice. Now, if we have a look at this one, this is kind of like a little test. The orbit test for a galaxy didn't work at all. So some things are broken on it completely. And then this was kind of like a cool Synthwave background as well. Then we actually created a crypt game, as you can see right here. And it looks pretty nice. I mean, again, it's smooth, easy to use, et cetera. Not bad at all.
But when we actually look at it on the benchmarks, you'll see how it compares in a second. You also might be wondering, okay, how does it compare versus GLM 5.2? I'll do a deeper tutorial on this later. But basically, if we have a look side by side, this is the maze example from Sonnet 5 Pretty nice. I do think it's a big step up versus what GLM 5.2 created over here, which is, you know, using the same prompt, but it's just a little bit buggy, not quite as smooth.
You can see, I mean, a little bit buggy. I would say very buggy at this point. So when we've compared it on tests like that, it didn't do so well. But then there were other things where, for example, the orbit didn't work at all with Sonnet 5 But if we have a look at GLM 5.2 version, it actually works really nicely and it can complete the task. So there were some tasks where Sonnet 5 actually totally failed and then GLM 5.2 crushed it. But we'll come on to more stuff like that later in a different tutorial. Now, we can also see here, for example, side by side, we've got two outputs from Sonnet 5 side by side for this game, as you can see. You can see it didn't do so well in terms of public reception.
So for example, you can see this tweet from Lisa and he said, Sonnet 5 goes straight into the garbage bin.
1.2x more expensive than Opus 4.8 Max, two times more expensive than GPT 5.5x High, five times more expensive than GLN 5.2. I mean, it's pretty wild when you put it like that.
I can understand why a lot of people are quite negative about it, because basically Opus 4.8 is outperforming Claude Sonnet 5 on many benchmarks as you saw earlier. But Sonnet 5 is more expensive than Opus 4.8. So Opus 4.8 is better, but Sonnet 5 is more expensive. How does that make sense? And again, I think with Fable 5 coming out soon, this is going to be one of those things where people don't even talk about Sonnet 5 anymore, because it's just it's been received so badly. You see another tweet here from Bridgemind. So he said, Anthropic seriously messed up with Claude Sonnet 5 The token efficiency is so bad, it actually ends up being more expensive than Opus 4.8. What are they doing?
5 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000775063198