Topics: Technology, Business, Investing
**Josh** (0:00)
Over the last 90 days, Frontier Labs shipped 15 plus models. OpenAI shipped three, Anthropic shipped four, Google shipped three, even Meta and Elon Musk shipped Frontier models. The Chinese shipped a bunch as well. And so if you're listening to this show, you're probably wondering, which model should I be using right now? Most of you likely have a GPT or Claude subscription, but you're wondering, should I be using these different models? The truth is, the Frontier landscape of AI models, when it was originally thought to be one or two, has now expanded to hundreds and hundreds of models. In OpenRATO alone, you can access 400 plus. And so on this episode, we're going to dig into which model you should use, for what particular use, and when makes the most sense. And you'll realize that the argument has shifted not from using the most intelligent model, but maybe using the most cheaper model, or using the model that's specific for you.
**Ejaaz** (0:48)
You got options, baby. There's a lot going on in the AI space, and we're going to help navigate that space, because there is a lot of options and it does get overwhelming. Between the two of us, we've probably touched every single one of these, at least a couple of times here and there. So yeah, between our 75 different subscriptions, we have covered these. We have some feedback about which one to use for when, which is best for which use cases. And I guess to start, we have this pretty helpful visual companion artifact here that can walk through what we're thinking, how we think about this. The first is the split. There are two distinct classes of model when it comes to considering which one to use. The first is open source. These are Chinese models predominantly. Most of these are your Kimis, your Deepseeks, like all of those models are Chinese open models. And then we have the actual US models that are all closed source. There's a big difference between the two. And we could see if we scroll down a little bit, the difference in usage between these two, because in June of 2025, US labs accounted for 70% of the tokens generated, and now they're down to 30%.
The economics have changed widely. So now that 30% is worth much more than 70%.
But it is interesting to note that there has been this kind of reinvigoration of Chinese models over the past couple of months. Now, granted, this is based on open router data. Open router is a single model router instance. This is not reflective of the norm, but it's just worth noting that some of these open source models are pretty powerful for the task at hand.
**Josh** (2:12)
Yeah, I think no one can debate the fact that people have shifted a lot using these open source models. And it's for a variety of different reasons. People want to own their own data. They want to run it privately at home. But the biggest shift, the biggest reason, has been because these models are a lot cheaper. And OpenRatadata, you mentioned, only really represents a very niche sector of software engineers that want to experiment with a bunch of these models. But even in the enterprise world, where OpenAI and Anthropic are pretty dominant with their own model share, they've been losing the token market share to enterprises that are trialing and testing different models to save tens to hundreds of millions of dollars as well. And when I look at the main reason why, it's not because the open models are the most intelligent. You'll see in a second as we go through the scorecard, GLM 5.3 from China, amazing model, not as good as Fable 5 You'll look at KimiK 2.7, you'll think the same thing. But those models are good enough to do the bulk of your own work. It doesn't sound too much like AI Slop. It actually just speaks to you normally. And the biggest advantage is it has fewer safeguards, which has its pros and cons, but it basically does the tasks that you ask it to. Now, overall, it's fine if we talk about a bunch of models on this show, but it's good to get an overview, essentially, of how intelligent or how effective these models are. And what we have on the screen here is something known as the Artificial Analysis Index or Intelligence Index. And this is basically the best benchmark to test the general intelligence of these different models. Now, it may come as no surprise to you, but the Claude models, Anthropics models, top this. We've got Opus 5 at the top, which is their most recent model launch. We've got Claude, Fable 5
28 more minutes of transcript below
Thousands of transcripts fetched by people building searchable podcast archives
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/YOUR_EPISODE_ID