Chinese AI Model Kimi K3 rivals top US AI models artwork

Chinese AI Model Kimi K3 rivals top US AI models

Elon Musk Podcast

July 18, 2026

The Chinese startup Moonshot AI has recently introduced Kimi K3, a massive 2.8-trillion-parameter model that represents a significant milestone for open-weight artificial intelligence.
Speakers: Jane, Paris Hilton
**SPEAKER_1** (0:00)
This episode is brought to you by Accenture. When you're advertising operations fall out of sync, everything else follows. Spotify and Accenture are working together to reinvent the rhythm of ad sales. Using automation, analytics, and smarter workflows to simplify campaign delivery and access better data across the business. The result? Less time spent on operations, more time connecting brands with the moments and fandoms that matter most. Learn more at accenture.com/spotify.

**SPEAKER_2** (0:30)
Right now, get up to 15% off select storage solutions. Put heavy-duty HDX totes to good use, protecting what's important to you. The solid impact-resistant design prevents cracking, and the clear basin sides make items easy to find even when the totes are stacked. Find select shelving and tote storage up to 15% off at the Home Depot to organize every room in your home from your garage to your attic.
Visit homedepot.com how doers get more done.

**SPEAKER_3** (1:00)
Are all batteries the same?
That's like asking if all soccer players are the same. Take Messi, the most decorated player ever. Is there any other player who has achieved that? No, just him. Now take Duracell. Is there any other battery with power-boost ingredients inside? No, just Duracell. Remember, goats only trust goats, because they're built different, and Messi only trusts Duracell.

**SPEAKER_4** (1:30)
Moonshot AI just dropped Kimi K3, which is an open weight model with 2.8 trillion parameters, and a 1 million token context window that directly rivals the absolute top proprietary models from OpenAI and Anthropic.

**SPEAKER_5** (1:47)
Yeah. I mean, you're looking at an open weight model coming out of a Chinese startup that is immediately challenging the absolute ceiling of US artificial intelligence capabilities.

**SPEAKER_4** (1:57)
Right. So with open models now matching these proprietary systems that we previously thought were untouchable, does choosing an AI model basically become a standard procurement decision instead of a technical one?

**SPEAKER_5** (2:09)
Well, before we even get to the procurement side of things, you really have to look at the raw numbers to understand how we actually arrived at this point.

**SPEAKER_4** (2:15)
Yeah. I mean, we have the current intelligence index scores right here, and the gap between open models and the proprietary flagships is, well, it's now measured in single digits.

**SPEAKER_5** (2:23)
Exactly, because you've got Fable 5 sitting at a 60

**SPEAKER_4** (2:25)
Right, and GPT-56 Sol is right there at 59, and then Kimi K3 is right behind them at 57, with GLM 5.2 following at 51

**SPEAKER_5** (2:34)
A two-point difference on an index like that is practically invisible for most enterprise use cases. Like if you're building an application to parse legal contracts, that core task is completed successfully by a 57 or a 60

**SPEAKER_4** (2:47)
Well, we also have the real-world testing data, specifically the GDPVAL AAV2 test, which measures agentic tasks against a human baseline. K3 scored a 1668

**SPEAKER_5** (2:58)
Which places it behind Fable 5 I think that scored 1760

**SPEAKER_4** (3:01)
Yeah, 1760 But K3 successfully leapfrogs Claude Opus 4.8, which scored 1600, and GLM 5.2, which is sitting around 1520

**SPEAKER_5** (3:11)
Right, but winning a single reasoning benchmark doesn't necessarily equal utility. We see models get optimized for these static evaluation harnesses constantly.

**SPEAKER_4** (3:20)
If you're not subscribed yet, take a second and hit follow on whatever podcast app you're using. It helps us keep making this. We appreciate you being here.

**SPEAKER_5** (3:27)
I mean, that's why we have to look beyond just the reasoning puzzles. The thing about K3 is that it consistently leads in practical tasks like automation, managing spreadsheets, web browsing.

**SPEAKER_4** (3:36)
Which are way better indicators of real world usefulness, honestly.

**SPEAKER_5** (3:40)
Exactly. When we talk about an agentic task in this context, we're talking about a process where you give the software a high-level goal and just let it figure out the intermediate steps.

**SPEAKER_4** (3:50)
Right. If you ask a standard model a question, it just generates an answer. Yeah.

**SPEAKER_5** (3:54)
But if you ask an agentic model to compile a list of competitors, it has to actively decide to open a browser, search for the companies, scrape the text.

**SPEAKER_4** (4:03)
Realize when a website is blocking it.

**SPEAKER_5** (4:05)
Yes, exactly. It has to back out, try a different search term, format the data and then present to you. So scoring 1668 against a human baseline means it is navigating all these multi-step failures and retries autonomously.

**SPEAKER_4** (4:20)
The spreadsheet automation aspect is particularly wild to me.

**SPEAKER_5** (4:24)
Oh, for sure. For a language model, a spreadsheet is this highly complex spatial environment. It's not just lines of text.

**SPEAKER_4** (4:31)
No, it's a strict grid system.

**SPEAKER_5** (4:32)
Right. Where the value of one cell is inherently tied to formulas in like three other cells on a completely different page of the document.

18 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000777313108