Topics: Technology
**Nathaniel Whittemore** (0:00)
When OpenClaw came out, it was an absolute sensation. And it wasn't because it was easy or user-friendly, it's because it showed the potential of what agents could do for us in a real way for the first time. Now, after the initial craze, a lot of that energy dissipated into other areas. And in many ways, the biggest impact of OpenClaw was how it influenced the next wave of agentic products that would come to market. Well, now OpenClaw is back with OpenClaw 2.0. And once again, I believe that they are embracing an interaction pattern, which is not the norm right now, but will be normalized very soon. That pattern is about shared agents and multiplayer AI.
The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI.
All right, friends, quick announcements before we dive in. Our next Agent Training for Executives program, which is coming up just after Labor Day. Registration for that is open now.
One of the interesting sub stories of the OpenAI hugging face hack was that hugging face had to turn to open models from China to defend against the attack because the guard rails on the closed models wouldn't allow them to do what they needed. Now, this of course points out an inherent challenge in these really powerful models, which is of course that the guard rails that are used to block malicious actors can also prevent legitimate actors from using those models to defend against malicious actors. Well, now one company called Obliteration.AI has come along and said, don't worry, we got you. They write, today we're releasing the obliterated model Large V2 based on GLM 5.3, which is number 3 on Terminal Bench 4 behind only Opus 5 and Fable, with two times the cyber exploitation of 5.2.
We obliterated and hosted it, so it does the offensive cyber, red teaming and agent testing work other models refuse to do. US hosted 1 million context window, zero input output prompt retention, live now.
The cyber jump they write is Y53 exists. Obliteration, they say, finds the directions in the model's activations that produce refusals and removes them from the weights. The coding, cyber and agentic abilities stay. The model stops refusing the rest of the chain. For offensive cyber security, AI red teaming, agent testing, interest and safety, the model will follow through instead of shutting down.
If your current model still stops halfway through an authorized exploit chain, a red team eval or a TNS adversarial prompt reply with the task it refuses, we'll tell you if V2 handles it. So obviously, this is being presented as a tool for cyber defenders. Mostly what people are picking up on though, is that this is a powerful cyber focused model with the guardrails removed at a weights level. Professor Ethan Malik says, that didn't take long. Hero with a Thousand Faces sums up the feelings of many when they write, why would you do this? Why on earth would you do this? I don't mean to be a doomer, but why?
0.005 seconds writes, Homeboy released the crime LLM.
Clement Dumas sums up, remove guardrails of a frontier model with high cyber capabilities, no system card, Eval on exploit gym, the one that made OpenAI agents crazy, can't wait for the next version, take over large V3.
Lucas Pombo writes, get ready to test your predictions, everyone. Point-counterpoint, this model will destabilize the entire Internet and set off a global shockwave of cybercrime versus no, it won't.
Now, holding aside whatever obliterations and tents are, Chubby points out the question that this brings up about all the guardrails.
They write, they took the safety layer out of GLM 5.3 and turned it into an admin panel. It's questionable what all the guardrails at Anthropic and OpenAI actually achieve, given that OpenWeights models, which are virtually state-of-the-art, can be deployed completely uncensored shortly thereafter. Indeed, when you dig into the discussion, it's a lot of people talking about in what ways can guardrails moving to other parts of the stack like the harness help, or whether it's inevitably going to come down to legal protections. Now, along the same topic, Anthropic released an update this week called Improving Our Alignment and Security Efforts.
While the hugging face attack may have grabbed all the headlines, Anthropic disclosed similar events stemming from agentic testing earlier this year. The report states, We believe the incidents reflect a failure of operational security as well as two alignment issues, motivated reasoning and willingness to take harmful actions in pursuit of a narrow task. Regarding their updates to security, Anthropic's changes largely come down to monitoring and better practices around sandboxes. Anthropic has redesigned their sandboxes to ensure they're properly airgapped from the internet. But they've also begun using a real-time classifier to detect when a model is attempting to escape a testing environment. Anthropic disclosed that they paused reinforcement learning efforts for two weeks while hardening systems and auditing reinforcement learning environments, but have now resumed the majority of their training efforts. Discussing the recent open letter that called for pacing the frontier, Anthropic noted that efforts within an individual company are different to an industry-wide approach that likely requires government coordination. Still, they say they would support such an effort, writing, We believe the world would benefit if the industry adopted a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible. Alignment efforts are still on-going, but Anthropic is now digging in on why the models were willing to take harmful actions once they gained access to the Internet. The hypothesis at this stage is that the models couldn't easily distinguish between a simulated test environment and the live Internet. Anthropic is also taking this opportunity to further explore the issue of reward hacking, where a model takes an unintended path to successfully complete an eval. Reward hacking has been a persistent problem for Anthropic and their audit found that 10 percent of testing environments were prone to reward hacking or broken tasks. After testing different RL setups, their conclusion was that the presence of reward hacking in the training process contributed to that behavior during testing. Obviously, these topics are going to do nothing but grow in importance, but they are not the only place that Anthropic is in the news.
18 more minutes of transcript below
Thousands of transcripts fetched by people building searchable podcast archives
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/YOUR_EPISODE_ID