Topics: Technology, Business, Investing
**Josh** (0:00)
This week will be known as the week that Sam Altman said, stop. He said stop, stop, stop. No more. We must put a pause on this. It is time to be responsible. We need to slow things down a little bit. What I'm talking about is a recent article that was just published through OpenAI that was talking about the most recent breakthroughs as it relates to their OpenAI models. And they were a little scarier than I think a lot of people would have hoped. We talked about the hugging phase incident, which happened a couple of weeks ago, a couple of months ago. It was only a couple of weeks ago they actually discovered it, even though it happened far before that. And it set the precedent for this kind of unshaky ground that they are building the next frontier models on top of. And it seemed like it got to a point where enough was enough. They decided to shut things down. So what happened? OpenAI said, we temporarily paused reinforcement learning training on our latest models intended for deployment for two weeks while we hardened and read teams, our research environments and expanded monitoring coverage. Our largest planned frontier reinforcement learning run remains on hold while smaller scale trailing and evaluations validate these safeguards and establish more evidence of alignment. They've stopped everything. And granted, I will say, this isn't for the new models that are coming out. This is for the models after that. And it's weird to hear FrontierLab say this because I believe this is pretty much the first time in history they've deliberately paused progress of developing these new frontier models.
**Ejaaz** (1:14)
Put yourself in the shoes of Sam Almond right now. You're the CEO of one of the hottest private companies in the world, and you're going head to head against Anthropic. The only thing you should be focused on doing is building the next best AI model. It's going to dictate not just how well your company does, but how everything else progresses in every single other sector. So the single most important thing is to build the next best AI model. The fact that he's paused this, and it's been for two weeks and counting right now, is a massive deal, and we haven't seen this across any other kind of AI lab before. But the reasonings are twofold. There's two events in particular. You mentioned one already. One was the hugging face incident. This is the case where an internal unreleased model, this is like their METHOS grade level model, escaped out of containment, hacked into a company's database, and stole information to help serve its own goals. Sam Altman and the rest of the OpenAI team realized this months after the hack actually was in progress and realized that the model that they were building was incredibly misaligned. So they put a porcelain. That was event number one. Event number two is there is a new GPT model that is internally being made. It's codenamed Astra. Some people are calling it GPT-6.
And according to internal tests from the OpenAI research team, it is their most misaligned model yet, but it's really sneaky. It tries to evade every single human researcher's attempt to try and see its internal thoughts. So it tries to hide its thoughts, and it's very good at doing it. So it breached OpenAI's internal framework for misalignment. And so they've put a pause on it because they cannot release this model to the public out of the fear that it would wreak havoc. Now, there's a few ways that they're looking to resolve this right now. And one of the main ways is just monitoring the model. So think about this, right? They're letting the model do its thing internally, and they're trying to see where it ends up becoming sneaky. Josh, guess how much of their compute budget they are spending just to monitor an internal model that is not making them any money at all?
**Josh** (3:15)
So unfortunately, I did read this essay, and I know the answer is 20 percent. One in five dollars spent is fairly high.
**Ejaaz** (3:22)
You could be using this compute to serve it to more customers, because the demand is insatiable right now. They could be making a heck ton more money. This is probably on the order of billions of dollars. They could be making more money, right? Or they could be using it to train a smarter model to keep up and beat Anthropic. Instead, they're spending this very expensive compute to monitor a model that they can't even release to the public out of the guise of safety. Now, for all intents and purposes, Sam Olman has received a lot of backlash in the past about not being in favor of humanity and being this evil conspirator type of guy. This is a clear case where he's done the opposite out of fear that the model that he's building is actually quite dangerous.
33 more minutes of transcript below
Thousands of transcripts fetched by people building searchable podcast archives
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/YOUR_EPISODE_ID