NEW GLM 5.2 BEATS Claude? artwork

NEW GLM 5.2 BEATS Claude?

AI News Today | Julian Goldie Podcast

June 16, 2026

GLM 52 vs Qwen 37 Max vs Claude Opus 48: Real-World Tests vs Benchmarks (No Second Chances)The episode compares GLM 52 (ZAI), Qwen 37 Max (Alibaba), and Claude Opus 48 (Anthropic) head-to-head on five one-shot tasks, arguing that benchmark rankings didn’t match real usability.
Speakers: Julian Goldie
**Julian Goldie** (0:00)
Glm 5.2 versus Qwen 3.7 versus Claude Opus 4.8. I put all three head to head, same five tasks, one shot each, no second chances, and the results flipped everything I thought I knew about picking an AI. So here's the part that gets me. The model that wins on paper, the best one, with the best scores came in last when I actually used it, and the one that looked the best, that actually shipped with no official scores at all. So if you've been sitting there trying to figure out which AI to actually use for your business, and you keep seeing these big benchmark charts to say, you know, something different every single time, well, we're going to look at what performs the best out of Qwen, GLM 5.2 and Claude Opus 4.8.
Now, these are three of the top AR models right now. GLM 5.2 is from a Chinese company called Jio Pu ZAI on this coding plan. Then you got Qwen 3.7 Max from Alibaba and Claude Opus 4.8 from Anthropic. Now, I ran all three through the same jobs, and we'll start this off. So the first job that we gave it here was a sort of voxel runner game. So we've got GLM 5.2 over here, Qwen 3.7, Opus 4.8.
So this is the one from Qwen 3.2, sorry, from GLM 5.2, as you can see here. It's a lot of fun, pretty interesting game, pretty cool, and a lot of fun to play. If we look at this one from Qwen 3.7, this is Qwen 3.7 Max, by the way, and you can see here that it's quite boring to play with, right? I mean, look at that, it's kind of buggy, but it does the job. It does the job. Then we have Claude Opus 4.8, and look how basic that is. That is not so much fun at all. So on the first coding test here, we can see clearly the Glm 5.2 is winning, and Qwen 3.7 comes in second. Next up, we have the inner system orbit map. And actually, if you look at all three of these, undeniably Opus 4.8 wins this, right? It's doing the best job here. And it's done something amazing, as you can see. And we can change the speed, we can change rotations, et cetera. This one is OK.
This one is pretty bad. Now you can like zoom in and it looks better and that sort of thing. On the outside, doesn't look that great. I mean, it's kind of cool to play with. But I think out of all of these, you know, Claude Opus 4.8, absolutely nailed it. Now we've got the liquid in a bowl test. So we have Qwen 3.7 over here, Glm 5.2 and Opus 4.8.
Now, if you look at these, I mean, it's kind of a boring test, but you can see here that the animation from Glm 5.2 is really nice, like we can change this, we can change the theme.
It's pretty cool to play with. We have a look at, for example, the one from Qwen 3.7 Max, not quite as fun. It just kind of fades out very quickly. Then if we have a look at the one from Opus 4.8, look how boring that is compared to what Glm 5.2 created. This is way better, way more fun, way more interesting. And that's what we want really.
So on test three and one, Glm 5.2 won and then on test number two, which is the Galaxy orbit, you can see the Opus 4.8 one. Then we've got the landing page test. So this is useful if you're checking, for example, you know, if it's actually creating something useful for business. So creating a website. Now let's have a look at this. This is Opus 4.8, super basic, super boring, not much to it at all. Not that interesting. If we have a look at Qwen 3.7, it's okay. I mean, this is kind of weird because there's nothing here, right? It's kind of just like an empty canvas, but the rest of it was okay. Then if we have a look at Glm 5.2, if we scroll down, it's got some nice animations. There's a lot more to the page, nice, nicely, but cleanly as well set up. And I like even like the animations of the page, it looks super nice. And you see how it's actually filled in the canvas, whereas for example, Qwen 3.7 didn't do anything. And the one from Opus 4.8, super boring. So that's the difference. This one, the landing page test as well. Then we have the arcade game. So this is pretty cool, pretty fun from Qwen 3.7. But the only issue is you see how the ball doesn't actually bounce off the walls, like it just disappears completely. If we have a look at Opus 4.8, this has built something better that is more playable and more useful. And then if we have a look at Glm 5.2 here, look how cool this is. This is way more fun and interesting. And so Glm 5.2 won on pretty much all of the tests there, apart from one, which is mind blowing in itself. And so in terms of the actual tests I've run here, I would say the Glm 5.2, which is a new model from ZEI, is actually beating Opus 4.8 and definitely beating Qwen 3.7 Max. Now, if we actually have a look at the benchmarks here, Qwen 3.7 Max is the strongest of all three. You know, Alibaba reports 80.4% on SW Bench Verified. We don't have the benchmarks. We only have Glm 5.1 benchmarks. So it's not really that useful. But if we were comparing side by side, Qwen versus Claude, well, it's actually beating Opus 4.7 on agentic coding. However, one thing to note here is the Opus 4.8 was so new when Qwen 3.7 Max came out, that they didn't include Opus 4.8 on their benchmarks as well.

5 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000773016991