Topics: Business
**Jordan Nanos** (0:00)
All right, Max, we're going to do a podcast. We're going to talk everything about Kimi K3 and maybe some other models that just came out. How are you doing?
**Doug O'Laughlin** (0:07)
Doing great. Looking forward to it, and thanks for having me, Jordan.
**Jordan Nanos** (0:11)
I'm not having you. Josh, thanks for having me. Okay.
On the docket, let's say Kimi K3 hot takes. Is this the third best model in the world? Impact on OpenAI, Anthropic, architecture changes, personal usage that we've had so far, what we think about their open source strategy, and maybe more. All right, Max, quick hot take. Is this the third best model in the world right now?
**Doug O'Laughlin** (0:36)
I think the answer is a clear yes.
People love shooting on benchmarks. I think benchmarks definitely have their problems. But I think sort of if you take a composite of all the main benchmarks and just look at the model rankings, they have been directionally correct over time. And I think if you look at that composite today, there's like a pretty clear top three with Fable, Sol 5.6 and now Kimi K3. And there's her just like always above everyone else, which includes of course other open source guys like, you know, DeepSea, CongealM, whoever. But also like very notably, it includes Google and Meta and SpaceX. So I think it's honestly it's an truly impressive and a very remarkable feat from the Moonshot guys.
Google in particular, I think should feel incredibly embarrassed right now that at one point guys, remember as as recent as like November, December 2025, everyone thought that like the clear AI dig three was Google, Anthropic and OpenAI. And even when I talk to like, you know, boomers today, they still seem to think that your top three is Google, Anthropic and OpenAI. And it's just like, clearly not the case anymore. So yeah, I'd say definitely third best mall in the world. I do think it's overall still worse than Fable and Soul 5.6. Kind of funny that they explicitly said that in their like model release blog post.
Maybe it's some like old fashioned Chinese humility. Maybe it's like they don't want to, you know, incur scrutiny from the US government or anything, because obviously there were some like delays for the Fable and 5.6 release. But overall, very impressed with them all.
**Jordan Nanos** (2:25)
Yeah, in the limitation sections of the blog post, they said, despite being a highly competitive model overall, K3 nonetheless exhibits a noticeable gap in user experience compared with Fable 5 and GPT 5.6 goal. So my experience using this personally is that it is good. It's really slow, which is really annoying. It's motivated me to try open source harnesses for the first time. And so I feel like I'm learning more about the harnesses than I am about the models, because frankly, all these models are like good enough to do the basic work that I've been doing so far. I can't really find a lot of complicated stuff that it can't do, which in and of itself is a bit of a feat.
Here's my hot take. For me, this might be the second best model in the world right now, because every time I try and do something meaningful with Fable, I get rejected and I get sent down to Opus. Even though I don't know if this is better than Opus, it is less annoying to not get rejected whenever I'm trying to do something. However, I'm not getting rejected when I use my API key and paper token, but I am hitting limits whenever I try and just use the web console or deep research or the coding plan. I haven't used the coding plan, but some other guys at SemiAnalysis have. And so it leads me to be like, what is the strategy here? Because these guys clearly just do not have enough GPUs to serve the demand, that they're up for this model. And previously, that was solved by an open-source strategy where they just drop the weights and then other people serve it and serve that demand. But they haven't dropped the weights yet. So I think they said, weights in 10 days or something?
What do you think the strategy is for the delay between announcement of the model, the API being available, and no weights yet?
**Doug O'Laughlin** (4:25)
Yeah, I mean, I think to be clear, this is all just pure speculation on my part. But I think one big reason is they need to give the VLM and SULing guys enough time to make sure they can serve this model performantly. Because if they just dropped it today, you have all this hype, but then everyone else serving the model is only giving you 20, 20 seconds or something. That's probably really bad for the brand, is it really capture.
22 more minutes of transcript below
Thousands of transcripts fetched by people building searchable podcast archives
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/YOUR_EPISODE_ID