**Julian Goldie** (0:00)
Claude Sonnet 5 vs GLM 5.2, The One Shot Showdown, Who Wins? We're gonna walk through it today, and side by side, we'll be comparing how Sonnet 5 compares to GLM 5.2. So let's get straight into this. And the first thing that we're gonna start with is a crypt game, like a dungeon crawler that we've created with both of these models. So this is GLM 5.2, this is Sonnet 5 Which one wins? Let's compare them side by side. We'll also compare the benchmarks in a second as well. So if we have a look, this is the output from GLM 5.2.
And not bad, not bad at all. Let's have a look at the output from Sonnet 5 I'd actually say the graphics are smoother, but the gameplay itself is not that great. Like there's literally nothing going on here. It's super dark and there is no game to be played. It's just you walking around a maze. So not that great. And I would say this is kind of the theme that I've seen across everything. And we'll come on some more tests in a second in terms of how they perform. So we've also got another one here, which is like, you know, build a ray caster maze. So let's see how they perform. This is the output from GLM 5.2. And if we have a look here, you can see that is super buggy, right? It's really, really buggy, as you can see here. And whereas, for example, if we have a look at the output from Claude Sonnet 5, it actually does win in this particular contest. So it is possible to get better outputs with Sonnet 5 than GLM 5.2. Bear in mind that Anthropix Sonnet 5 is closed, whereas GLM 5.2 is open source. So that plays a big difference as well. And also, of course, GLM 5.2 is a lot cheaper than Sonnet 5
So let's see how they perform on the test here. If we compare them on CursorBench, we have Fable 5 max right at the top, extra high. So Fable 5 is absolutely dominating, crushing here. If we have a look at Sonnet 5, it's way down the charts.
You can see here, for example, it's number 13 However, GLM 5.2 is actually a lot lower. So you can see it scores 61.2% for Sonnet 5 on CursorBench, whereas GLM 5.2 scores 54.6. So there is a big difference here in terms of how they perform side by side. Now you also might be wondering, okay, should you pick Opus 4.8 or should you pick Sonnet 5 if you're using Claude, Claude Code, et cetera? So what we can see so far is the Opus 4.8 crushes Sonnet 5 on the benchmarks as well. So if you have a choice between them, I'd go Opus 4.8. But bear in mind that Fable 5 is expected to drop within the next 24 hours. That's been announced officially from Anthropic. So you know, if you had a choice, I'd put Sonnet 5 down here. Then we've got Opus 4.8 above that. And then we have Fable 5 that's even better than all of them. And also something to bear in mind as well, you see in this tweet by Lisa, one point is 1.2 times more expensive than Opus 4.8. And Sonnet 5 is two times more expensive than GPT 5.5. But actually five times more expensive than GLM 5.2. So bear in mind like if you're using the coding plans, so you can get a coding plan with Claude, or you can get a coding plan with GLM 5.2.
If you're using the coding plan, it's going to be a lot cheaper with GLM 5.2. And also bear in mind like if you want like the full power of GLM 5.2, but you prefer using something like Claude Code, we actually set up something called GLM Code inside the Agent OS system, so that we can use the power of GLM 5.2's brain inside Claude Code. And then we can say everything that we've created right here. And it performs pretty well when we're doing that as well. So you get like the agent harness of Claude Code, but you can plug in GLM 5.2 and then, you know, either way, whatever drops, whatever comes out next, you can get the most of these models by having great systems like this. Now, let's have a look how it scores on Goldie Bench. So obviously Claude Sonnet 5 just dropped a few hours ago.
It's a million-token context window. It's scoring pretty well on SW Bench Verified, to be fair. They've actually said it's their most agentic Sonnet yet. But, I mean, that should be the case anyway, right? I would expect that anyway. Like, for example, if you're using the older version of Sonnet 5, are they going to be as good as the new Sonnet 5? No.
7 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000775082480