**Mikkel Rosenvold** (0:00)
Summertime is maybe good, summertime is maybe shit. Hello out there, welcome to Real Vision, welcome to Macro Mondays. My name is Mikkel Rosenvold, I'm your usual host, and I'm joined by the returning Andreas Steno, back from vacation in France. How are you, Andreas?
**Andreas Steno** (0:36)
Good, it's nice to be back. The temperatures here fit my weight a little bit better than we just put it like that.
**Mikkel Rosenvold** (0:46)
Yeah, and maybe the momentum in markets is well, Andreas, looking decently this morning, Andreas. Let's dive straight into the topic on everyone's agenda this weekend, the Kimi scare.
I want to get your take on it, you wrote an entire piece on it this morning. Has China taken over the leadership in the AR race?
**Andreas Steno** (1:06)
No, next question.
Of course, I think it's the first time that we've seen China leading the race in one subcategory in the arena AI scoring methodology, but they don't lead the race overall. But they have taken the lead in a subcategory related to front-end coding, which, I guess, is something that we need to consider, especially when we look at how markets should respond to this, because I'm not necessarily sure that I agree with the market takeaways from, say, Thursday, Friday last week. And I basically think the market got this wrong, but we can spend a few minutes on that, Mikkel.
**Mikkel Rosenvold** (1:52)
Absolutely, Andreas. So we had these redisco results that Kimi K3, the new moonshot model, was essentially the best out there. This caused a whole wave of panic about China taking over the AI leadership, driving this, and a lot of confusion, as I saw it on social media and elsewhere, around how do you even play, what does that even mean if China takes over?
Who's the losers? Who's the winners from that? So first of all, maybe let's dive, I know you like to dive a little bit into the actual fabric of this.
Is Kimi the best model? How much does this matter?
**Andreas Steno** (2:28)
So my preferred third-party provider of, say, artificial intelligence or scoring of artificial intelligence, is the provider called Artificial Analysis. I think we have a chart on the scoreboard from that on page 8
Again, yes, if you look at the front-end coding from Arena AI, which is, by the way, an old spin-out of UC Berkeley, the Kimi K3 model is ahead of Claude Fabel, for example. But if you look at it on text, which is basically the broadest use case for many, it is still behind, which is what you would expect. And if you look at it, broadly speaking, from this artificial analysis intelligence index, it is also behind both Claude Fabel and ChatGPT 5.6. So are they ahead? Well, in a niche corner, yes. I guess front-end coding is not the main use case for most users of AI, but it's of course an important use case. And top-down, I think this is the biggest issue for Anthropic.
If you look at the average user of Anthropic, they're probably more prone to use Claude Fabel for front-end coding purposes than the average user of ChatGPT, just to put it in very simple layman terms, right? So in a niche use case, yes, there is an issue here potentially. In a broader perspective, I still see a lot of issues with the notion that China is ahead.
Let me just, again, paint a very simple picture on page 9, Mikkel, because what I did was that I looked at the overall scoring of the best Chinese model relative to the best US model on all scoring parameters, right? The artificial analysis index again. Nothing has changed, basically. What we typically see is that a new flagship model is shipped by Anthropic or OpenAI. The gap opens again, so China is, say, 5-10 percentage points behind. Then we get a new flagship model from China, catching up, but never getting past the frontier parity goalpost. It's essentially the same that has happened here. The true shocker was when Deepsea got to about 95% of the capabilities of the flagship models in the West in early 2025 They've never really gotten anywhere closer than this, say, 5-7 percentage points off the frontier model. I think there's a really, really important conclusion hidden in this chart. They're always behind. I think there's a reason why they're always behind, because they obviously use the frontier models to train their models. That would be my best assumption watching this. We can get into the weeds of a long technical discussion here. That's not by any means by my intention. The point here is that they're always slightly behind. I actually consider this a pretty fruitful setup in many ways, because if you have someone chasing you, like one month, two months after, it keeps you very attentive. It keeps you very focused on shipping the next model, etc.
21 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000777595000