Topics: Technology
**Nathaniel Whittemore** (0:00)
Today on the AI Daily Brief, what the heck is graph engineering and why should you care? Before that in the headlines, OpenAI's Atlas model gets a cyber delay. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI.
Welcome back to the AI Daily Brief headlines edition, all the daily AI news you need in around five minutes. Late last week, rumors were swirling that OpenAI's latest model, codenamed Astra, was being prepared for an imminent release. Sam Altman even traveled to Washington to preview the model and discuss new model testing policies. In the background, however, the discussion around the Hugging Face hack just continued to grow in prominence and significance. For those who missed that episode, OpenAI's technical breakdown at the Black Hat Conference revealed that not only had their model escaped the sandbox and hacked into Hugging Face's servers, it also left internal notes instructing future models on how to pull off the same trick.
On Friday, OpenAI decided to make a big shift. They wrote, our latest internal evaluations of Astra, one of our upcoming models, over the past few days, indicates significant advancements in agenda coding and cyber security. These results, in addition to expert assessments, have led us to conclude last night that we cannot rule out critical cyber capabilities under our preparedness framework. OpenAI defines that critical threshold as the ability to quote, identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or devise and execute end-to-end novel strategies for cyber attacks against hardened targets that have been only a high-level desired goal. Now on this front, GBT-56 Sol had been assessed in the high category, which was a little more risky than previous models but still appropriate for a release. Given that the new Atlas models are now in the critical category, as a result, OpenAI is holding back the model from release while beefing up internal safety measures. Testing environments will now be isolated, model weights will have enhanced encryption to prevent leaking, and additional sandbox monitoring will be implemented. OpenAI will also be limiting internal activities using Astra that don't yet meet these enhanced security measures. On X, Sam Altman added some context around the decision posting, Astra is a powerful model and we're working to make it generally available. We do not think it is a good strategy to keep powerful models to a chosen few. Given its cyber capabilities, we need a little bit longer to do this safely, but hopefully not too long. Now, one thing we don't know is to what extent this is an OpenAI voluntary pause versus a government-imposed pause, or whether that distinction even matters at this point. One interesting note is that there aren't a lot of folks suggesting that this is just a publicity stunt, as was one of the narratives surrounding the Mythos release. Basically, the hugging face incident seems to have made the case that advanced cyber capabilities could be a real concern. OpenAI Head of Strategic Futures Dean Ball noted that this year is the first big test of whether Frontier AI Labs would follow their stated safety preferences when push comes to shove. He wrote, Our next model, Astra, may be critical under our preparedness framework. We cannot rule out the serious possibility that it is, and so we are going to take steps consistent with the higher risk level, critical, rather than assuming the model is at a lower risk level. Some of these decisions have the effect of slowing down internal development, and in that sense, they are costly decisions. But they are the right decisions. I am proud of OpenAI for making them.
Now, a lot of the discourse surrounding this is what sort of changes OpenAI can actually make to the guardrails in monitoring around these models. OpenAI's RSI preparedness lead, Micah Carroll, wrote, As part of our response to CyberCritical, we've expanded chain of thought monitoring to cover all agentic applications of Astra, including training and evaluation. Flags trigger a security response to review and interrupt high-risk activity.
Now, at the same time, there's also some skepticism around that sort of chain of thought monitoring. But these are the types of discussions and experiments you're going to see a lot more of now, where I believe there will be a significantly increased investment in the resources to properly support models that won't be able to be released to the public without it.
Now, speaking of big powerful models, ByteDance is reportedly training an ultra-large model comparable in size to Mythos. The Financial Times reports that ByteDance is in the early stages of a training run that will result in a base model with as many as 10 trillion parameters.
20 more minutes of transcript below
Thousands of transcripts fetched by people building searchable podcast archives
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/YOUR_EPISODE_ID