**Akash Pasricha** (0:13)
Welcome, everyone, to The Informations TITV. My name is Akash Pasricha. It is Friday, July 17th. It's been a great week. We had our one-year anniversary party. We had our UBS shoot. We are back in our New York studio next week. But today on the show, Moonshot AI and Kimi K3 is getting a lot of attention as it challenges closed-source models. We'll bring on the CEO of benchmarking company, Arena, for the latest on what he's seeing.
We'll then unpack Netflix's latest quarter and separately, some other big headlines from the week on this week's edition of The Editor's Cut. We'll dive into our latest exclusive reporting on the traction that Google is getting with its TPUs. We'll close out the show with a conversation with the CEO of Beehiiv as they make a big bet on social. It's going to be a great show, so let's get right on into it.
Kimi K3 is the big news of the last 24 hours. Moonshot AI released the model. People are suggesting it could be a big threat to existing models dominance and an open source threat at that. I want to bring on Anastasios Angelopoulos, the CEO and co-founder of Arena for all of his thoughts. Anastasios, welcome back to the show. It's great to have you here.
**Anastasios Angelopoulos** (1:26)
Thanks for having me.
**Akash Pasricha** (1:28)
New day, new model.
**Anastasios Angelopoulos** (1:29)
New day, new model.
**Akash Pasricha** (1:31)
You've been busy, man. I mean, gosh, this is like the fastest pace of model releases I think we've seen in a while.
**Anastasios Angelopoulos** (1:39)
It's incredible. Six models in the last seven days. It's been Grok 4.5, Kimi K3.
Previously, we had Jalen 5.2. It's the new series of models, the GPT 5.6s, the Soul, the Luna, the Terra. It's been an extraordinary piece of model development, and it's left us with a Chinese open-source model on the top of the coterie.
**Akash Pasricha** (2:03)
Did it come out of nowhere here? Were we expecting? We'll talk about the reviews in a second. Were we expecting this to be this good initially upon reviews?
**Anastasios Angelopoulos** (2:11)
Well, there have been murmurs of the great performance of Kimi K3 for quite a while. But I'm not sure that people expected it necessarily from Moonshot. I will tell you what has been the trend, which is that over the past year, you've seen open-source and closed-source models had a big gap, and then it started closing, closing, closing. Then they were tracking with open-source models just a few months or weeks behind closed-source models. People were saying, they're never going to crack it, and it's because of distillation. They're just distilling, distilling, distilling.
Fundamentally, when you distill a model, it degrades and so you're never going to get better performance. But this is the first time we're really seeing the narrative that breaks that mental model, that says that the Chinese labs might actually just be really good at developing models and not just distilling American intelligence.
**Akash Pasricha** (3:03)
Right. So you run one of the most popular platform of benchmarks that is widely cited in AI. What is your data telling you about how good Kimi K3 is?
**Anastasios Angelopoulos** (3:17)
Well, the way our platform works is that we have a user base of tens of millions of people that are coming on arena to use AI for their real workflows. Their agentic workflows, they could do single threaded conversations that are tens or hundreds of turns long. They're getting to do workflow automation, they're doing coding, they're doing math, they're doing all these economically valuable tasks, and we take all those tasks and turn them into a benchmark. One of the benchmarks that we released for Kimi K3 has been code arena, and specifically, Kimi K3 is at the top, it's number one in front-end code arena. So it's the best at the lovable vibe coding use case. Now the best model in the world for that, of course, to which hundreds of millions of dollars in inference-suspended will accrue is Kimi K3.
**Akash Pasricha** (4:03)
So ahead of 5.6, ahead of Fable too?
**Anastasios Angelopoulos** (4:08)
Ahead of Fable, which is absolutely remarkable.
**Akash Pasricha** (4:12)
So let's just pause here for a second.
This is kind of, you know, Fable was the model that really caused a bit of a reckoning in terms of cost. This is an open-source model which, I mean, means it's significantly cheaper, right? Am I correct there?
**Anastasios Angelopoulos** (4:30)
That's correct. And I will tell you, the gap between Kimi K3 and Fable is not small. It's a win rate of in the order of tens of percent. You know, let's say 10% is the difference in win rate between Kimi K3 and Fable 5, which is, you know, it's not the largest gap that we've seen on arena, but it is substantial and indicates a real delta in the change in performance. And now that comes at the cost of a model that is a sonnet level cost.
46 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000777258872