Are Cheaper AI Models Better than Claude and ChatGPT? artwork

Are Cheaper AI Models Better than Claude and ChatGPT?

Limitless: An AI Podcast

July 16, 2026

We're discussing new AI model releases from xAI, Meta, OpenAI, and Anthropic, and the shift toward cheaper, more efficient models.Focusing on Grok 4.5 and Meta’s MuseSpark 1.1, there's also a broader move toward model routing and enterprise use.
Speakers: Josh, Ejaaz
**Josh** (0:00)
Just last week, in the span of about 48 hours, three of the most powerful AI labs on the planet all shipped brand new models. Elon shipped Grok 4.5, OpenAI took 5.6 global, and Meta, for the very first time in history, put a price tag on its own frontier model. And here's where it gets interesting. For the last five years or so, the deal in AI was that every time a model shipped, it got smarter and cheaper at the same time. But that kind of died this year with the Anthropix Fable 5 and these new frontier launches like GPT 5.6. So today, we're asking the question that probably everyone should, when frontier intelligence costs less than a cup of coffee, when they cost just a few pennies per millions of tokens, what does that look like in terms of your costs and how much you use these models? I mean, this is a totally different paradigm now. We have a very clear separation between Fable and 5.6 Sol, and Grok 4.5 and Meta's new model, MuseSpark 1.1. And it seems like these models are kind of diverging in a way that's really interesting, particularly centered around price.

**Ejaaz** (0:56)
For the last couple of years, the ultimate validation of whether your AI model is good or not is how intelligent it is, how smart it is, and no one really cared about cost.
And then Fable 5 kind of came on the scene, and I think it was like $20 or $30 input and like $80 output. And companies that were spending tens to hundreds of millions of dollars started to think, is this right? Does this make sense? Do I need the smartest model to do every single task? And so we hear the likes of XAI, Elon's company. We look at the likes of Meta. They released these models and we compare it to Fable 5 And we're like, they're not actually that smart. But that's the whole purpose of the models that they're releasing. They're not trying to be as smart as Fable. In fact, they're betting on the opposite. They're betting that the cheapest model per unit intelligence is the model that will ultimately fill the middle ground. That will be the ultimate model that is embedded in every single enterprise and used by every single user. Because the truth is, you don't need the most intelligent model to do your task. Maybe if it is 90% to 95% of the intelligence of the smartest model ever, but it costs one tenth of the price, it's a no-brainer that you're using these models at scale. And like you said, Josh, over the last week, there have been two particular companies, Meta, who we've known and spoken about a lot on this show, who have spent upwards of, I think, $35 billion to try and build the world's best model, came out with their new Meta, MuseSpark 1.1, and then you had SpaceX AI, who recently IPO-ed and there's a lot of pressure riding on them, released their new Grok 4.5 models. Now, are they as good as Fable? No. But are they cheap enough to use at scale when you're using or spinning up agents or when you're trying to figure out that long complex task that's going to take tens of hours and you don't want to burn very expensive Fable tokens? These models might be the ones to choose, and I think it's probably worth covering a bunch of them.

**Josh** (2:48)
Yeah, there's a new Meta almost in town, where it's like a new model doesn't necessarily mean higher intelligence. It could just mean higher efficiency. And I think that's where we're going to start with the Grok 4.5 release, because this is kind of like an opus class claim, but at a third of the price, which is a pretty big deal. So this new Grok 4.5 model, it's built on XAIs or I guess SpaceX AIs, their version 9 foundation model, which is about 1.5 trillion parameters. And for reference, the version we've been using all year, if you've used Grok at all this year, that is the version 8 small, which is about 500 billion parameters. We're looking at about 3 times multiple in parameter count, which generally speaking is 3 times better, but probably a little bit more. It seems like this is a serious increase relative to what we've been using. And Elon has called this Opus class model. But instead of just being this highly expensive, very slow model, it's much more quick and it's much more efficient. When it comes to how many tokens you're able to generate for that same dollar, and this is very clearly the route that you can see SpaceX AI has been trying to go for a long time. They're very hardcore engineers. They love the engineering challenge. And what they're trying to do now is figure out how you can kind of sculpt these GPUs that they're training on to get as efficient as possible. We spoke a few weeks ago, Ejaaz, about the Esht guys, like the startup who's building their own training architecture chip stack, and they're basically building their own servers. And within that, they're pretty big to make one specific thing work, which is the transformer. And we talked about how GPUs are not very efficient. They don't actually use, like, sometimes up to 60% of the GPU isn't used. What Grok and the SpaceX AI team are doing is they are taking that code base, and they're really getting down to the bare metal to figure out how to squeeze the most juice out of it. And that's what this model is. That's what 4.5 is. It has $2 in, $6 out per million tokens generated. And it seems like it's incredibly efficient. When comparing it to other models like Opus, it appears as if it's up to four point times fewer tokens needed in order to reach the same task. So that's like an adjusted multiple of what it says 17 times less than Opus. That's like a really big deal for a model as it relates to cost at least.

26 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000777066675