Claude Opus 5 review: this model is brilliant (but annoying) artwork

Claude Opus 5 review: this model is brilliant (but annoying)

How I AI

July 24, 2026

I’m tired of new models. Every week there’s a new benchmark, a new frontier intelligence claim, a new thing to test. But here we are, because Opus 5 just dropped and I’ve had real hands-on time with it, so you’re getting the honest version.
Speakers: Claire Vo
**Claire Vo** (0:00)
You guys, I'm tired. What I'm tired of is models coming out every week, new models, new benchmarks, new frontier intelligence, new things to test. It's been a little bit of a run in the past month. We've seen Fable come and go and come again. We've seen GPT 5.6, we've seen Sonnet 5, so many 5s recently, and just so many models. And I've been lucky. I've been able to test these models, been able to play with them for sometimes days, sometimes weeks. It just depends on who I'm working with. And it's been really interesting and exciting to have access to all this frontier intelligence. But I think we have an intelligence overhang. I really think that we're running out of, by we, I mean the average coder, average software engineer, average creator, average builder, average consumer, average business person. I think we're running out of ways to truly leverage this incremental intelligence. So this is my hypothesis in the next year. We're always talking a lot more about speed, talking more about cost, we're talking more about open source, and we're going to be talking a little less about intelligence, although I think we might be talking about specific types of intelligence other than software engineering.
But despite being tired, today we are going to talk about Opus 5, baby. Opus 5 is here, so we got 0.2 additional Opus points, Opus opals, whatever, however, we're tracking the increments here on Opus. Opus 5 is here. I've been able to test it a little bit. I have some opinions. Now, some of the stuff that I cover this episode is going to be a little different than what I've done in the past. Yes, we're going to do the How I AI benchmark live. And yes, we are going to look at the prototypes, we're going to look at PRDs, and we're going to look at agent personality. But I'm also going to put on my large language model psychologist hat, and we're going to talk about Opus' personality. And we're going to talk about Opus' personality relative to GPT's personality, because I think this is super interesting. If you're thinking about what is the difference really between these models, and you don't want to look at the difference in terms of benchmark capability, you really want to understand what these labs are going for, why these models are being built, and how they're being tuned. Looking at their personality at this moment, where intelligence is very high, is super fun. So we're going to do a little of that. We're going to do the How I AI benchmark. We might do some live coding.
We're not going to cover too much of the specs in the model because, read the blog post, read the blog post. We'll link to it in the show notes. What we're really going to talk about is, is Opus 5 good? Am I going to swap it in and how is it different than the other frontier models on the market? So let's get to it. First, let's just get it out of the way. Is Opus 5 good? Yes, it's good. Is it going to be all the benchmarks? Of course, it's amazing at benchmarks. Can it write code? Of course, it can write code. What did I test it on that really gave me a sense of its personality, which at this point where I could just simply cannot absorb any more intelligence, I really zeroed it on. And you know what? I haven't seen this since I would say Gemini 2.5.
This model is neurotic AF. It is so timid. It is so apologetic. It is so scared. I have never experienced this or I haven't seen this sort of neuroticism in a model in a while. And it's really funny. It bubbled up in a couple of ways. And I want to show you a few examples. Okay, let me just give an example of its timidity. And this chat was very long. There were so many examples of this where it was like, I think this is the answer, but do you think I should do it or do you want to do it or should we ask someone else to do it? It was like every time I just kept saying like, why don't you solve this? Why don't you do this? And this was a really good example. I pulled a branch and I was like, there is truly like a one line merge conflict. I could have not been lazy and literally just done this manually. I don't know, I was just feeling lazy. It was late at night, whatever. Like, can you fix this merge conflict?
And it was like, oh, but that's someone else's branch. Like, that's not my branch. I don't want to do that without him knowing. It's his commits. And if he has local work and flight, it might be disruptive. And I'm like, just do it, man. Just go. Like, go ahead. And this was like my constant experience with Opus 5, is it was like so, so, so timid. And so I just consistently had to say over and over again, like, man, just do it, make a decision. And then there was this really funny example when I spun off some subagents to kind of like assess the correctness of this query that we change from kind of like an ORM query to a SQL query. And it asked for things that it wanted a human on. It was like, can a human please check this stuff? Like, can it check this four megabyte ceiling? And can it check TypeScript and SQL? And can he like check for me? Because no one has confirmed this for me. And I was like, who is nobody? You're nobody. You said this sentence, nobody could confirm it. Like, can you just try? And then it went on the web and tried. And so it just has this like really interesting conservatism, neuroticism, human reliance that I think is super fascinating. And this gave me this inspiration to do something a little bit different this episode, which is I was like, I was going to go interview this model and figure out what is going on in its brain. Like, I'm going to figure out what it thinks about our relationship, because I just totally noticed this dynamic that I hadn't noticed in other models. And I hadn't really been attuned to before, where it was like very reliant on me as a human. And I'm like, I want you to be autonomous. And sometimes when I say go run subagent stuff, it'd be autonomous, but it wouldn't make decisions. And I hadn't seen a model delegate code to me in a really long time. And I was like, why are you asking me to write code, man? I only have 10 fingers. And so what I did, whether or not you think this is scientific or not, this is Claire's email, is I just went to the model. I went to Opus and I said, yo, who's smarter, you or me? And it gave me this like very anthropic-y answer, which is like, it depends what you're asking for. I could do these things better, but you can like feel if something feels wrong. And you can, this one was like so fascinating. It's like, you can tell which of your teammates is quietly burning out. I'm like, bro, Claude, I'm gonna burn you out. We don't burn out. The humans don't burn out on the chat PRD team. We burn out our agents, sorry, agents.

15 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000778228245