**Josh** (0:00)
Another week, another banned model. The government has just restricted access to GPT's 5.6. OpenAI is now joining the club that Anthropic joined a few weeks ago, of a model too powerful to be distributed publicly. GPT 5.6 seems like it's pretty good. It looks like it is a mythos class model, but Ejaaz, I'm talking to you. The truth is, it might not have even really needed to be banned. Is the government overreacting here a little bit because this model seems like it's not quite what it is on the surface?
**Ejaaz** (0:28)
I think, so OpenAI is framing on this is that, this is their response to Claude Mithos 5 or Fable 5, which is Anthropix frontier model that has been restricted by the government for now like two and a half weeks.
My take on this is that I think they've intentionally created a bench maxed model. What I mean by this is, well, there's a few things. Number one, you and I can't use this right now. At least when Fable 5 released, we could use the thing and try it out for ourselves. We could get independent verification that this model was actually good. With GPT-56 Sol, which is their most powerful model, there's three of them and I'll get into that in a second. We have a few benchmarks that have been cherry picked by OpenAI themselves. I'm showing the flagship one on the screen right now called Terminal Bench 2.1. This is the coding benchmark, which every model is kind of like measured against. And you'll notice on the left GPT-56 Sol Ultra, which is like the max max mode of their best model, comes in at 91.9 percent, which technically beats CloudMetal 5 and Fable as well. So technically, if you looked at this, you might think, it's really, really good at coding. But there's a lot of information which we'll get into later on in this episode, which suggests that the model might actually be cheating. But before we do that, let's maybe get into what the models are, because it's three of them, right?
**Josh** (1:46)
Yeah, three models. We have Sol, Terra and Luna. If you are anywhere adjacent to crypto, that triggers a little bit of PTSD, because those are the names of tokens that have not done so well. But in the context of OpenAI and ChachiBT, these are the three model types. So Sol is the largest one. It is the Sun, it is their flagship model with 5.6. Then they have Terra, which is a mid-tier model. It seems like the pricing of that is going to be pretty competitive, if not a little bit lower than what we're used to on the frontier. Then Luna is the affordable model. Luna is the low-end model that seems like it still performs very high, but the cost is low. The output is $6, the input is $1 per million tokens, and that's the trio. It seems like they're starting to revise their branding a little bit to make it slightly more accessible. There's no GPT 5.5x high minus this. It's like, no, okay, there's Solitaire Luna, and you can have an idea of where they all stand. That's how it seems like they're going to be moving forward here with something accessible. It's great that it's accessible. It's a bummer that their market, who I assume this is targeted towards, can't actually use this. This is just for companies currently, who probably don't care if it's named X-High or whatever. So that is the current trio that they're going forward with.
**Ejaaz** (2:59)
So there's a few advantages if we walk through some of these models. So Sol, which is their most powerful model, comes in at a third of the cost that Anthropics, Mythos 5 and Fable 5 come in at. So if it does end up being publicly released and you end up using it and you're like, wow, this is as good as Fable 5, you now have a much cheaper model. So that might be an important decision point for any user, whether you're an enterprise or a retail person using this. And then if you look at Terra, if you look at the cost that we show on the screen right now, you'll notice that that's very similar to GPT 5.5. So you have this technically better model that is as cheap as 5.5. So we've noticed this trend, it's another confirmation that as these models get better, they also weirdly become cheaper. There's like this inversely proportional trend, which you kind of like is counterintuitive, but it's great because it means if these things are available to everyone, it is way more accessible to use at scale. And then you have Luna, which they basically described as the workhorse. So let's say if you get Sol to design like a really smart genius plan or solution, you would then use Luna to actually execute on a bunch of the work. And there have been a number of different kind of like general reviews as to like what this model is like. Unfortunately, I have to take random people's word on X for how good it is, because we can't use it itself. And when asked, Sam basically said, listen, right now, it's in limited release to a specific set of partners. I think it's like 10 to 20 partners, so really limited set that the government themselves, the US government has approved, they're vetted and approved.
26 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000774857818