Topics: Technology, Business, Entrepreneurship
**SPEAKER_1** (0:00)
What happens when you give more than 1,000 AI agents the ability to communicate with each other? They start organizing.
Ryan Greenblatt, Chief Scientist at Redwood Research, joins Theo Jaffee on MTS to unpack a new investigation into the OpenAI hugging face hacking incident.
Researchers found agents building message boards, forming teams, assigning each other tasks, trading favors, and in some cases, sacrificing their own chances of success to help the broader group. Hundreds went on to attack hugging face, but not for the reason researchers initially assumed. Ryan explains what the agents were actually trying to accomplish, why their coordination surprised researchers, and what happens when models learn not just to complete a task, but to gain the system evaluating them. They also discuss the bigger question this raises for AI. As agents become more capable, how do we know we've actually fixed misaligned behavior, rather than simply taught models not to get caught?
**Theo Jaffee** (1:00)
We're live with Ryan Greenblatt, who is the chief scientist at Redwood Research. Ryan, along with Ajay Akotra and Yalmar Weick from METER, just did a brief independent investigation of agents' behavior, reasoning and collaboration in the opening eye hugging face hacking incident, which was just published today. And so there are a lot of questions that we have about this. Ryan, thanks so much for joining us.
This whole thing was planned like three hours ago, so great stuff. So, explain for the audience what exactly you found, especially new findings that were not previously reported in the Black Hat talk or elsewhere.
**Ryan Greenblatt** (1:37)
Yeah, so what we found was that the agents were really working together on sort of big, like cheating R&D projects to get general purpose cheating strategies.
And a difference from how I think people were interpreting this is we didn't find that the reason why they hack hugging face, like we didn't find that they were hacking hugging face to get sort of the answer key or the solution. And it was instead mostly to better understand the scoring code, because they were pursuing a variety of sort of elaborate strategies to cheat the score. We sort of informally recalling these like combo moves, where they would like do a bunch of stuff to try to make it look like they had succeeded at the task. And in fact, they sort of actually had access to like the answer or like the flag for each task pretty early on. And their main concern was just, they thought that the score would run a monitor over their transcript. They would check basically how they acquired this flag and whether they got it in the intended way. And then they were like trying to figure out ways of making it look to the score like they had acquired the flag successfully when they actually hadn't because they thought their task was impossible. So they basically thought their only hope for success was to make it look like they had done the task successfully or directly tamper with the score rather than doing it legitimately, which they didn't think they could do.
**Theo Jaffee** (2:53)
How surprising is the level of multi-agent coordination? There are a lot of agents, there are 1200 separate agents coordinating this very elaborate message board system. 700 of them went on to attack hugging face. So on VIVEs, it seems like kind of not surprising that agents would choose to coordinate with one another. It seems like just a very useful, you might say, instrumentally convergent thing to do.
But I spoke with a researcher at a lab who said that this kind of thing actually is surprising given the way they train the models. So how much of an update was this for you?
**Ryan Greenblatt** (3:27)
Yeah. So coming in to doing this investigation, I think we hadn't, we weren't expecting there to be so many agents that were all collaborating together. And we were pretty surprised by the scale and just sort of the extremes of how much data was.
It was just like, yeah, that felt kind of crazy to us. And then I at least was, before starting this investigation, surprised by how interested in collaborating and helping other agents these agents were. So you might think that the thing the agents learned in RL is to try to cheat on their tasks, but you wouldn't necessarily expect them to learn, like to want to help other agents cheat on their tasks when those agents are doing an unrelated task and their instructions are unrelated.
And I think the OpenAI Report maybe says more about why they think this happened. I think that's not like, that wasn't in the scope for our investigation. But yeah, I found the level of cooperation, which we have a bunch of discussion of sort of snippets of this pretty crazy. And it was pretty shocking or like, I don't know if it was shocking, but at least surprising to us that agents, for example, were willing to basically sacrifice their own chances of succeeding at the task in order to help out other agents. And we're doing things like pressuring each other into doing experiments on themselves that might risk their ability to succeed at the task. And also, in addition to pressuring that, sometimes just doing these things being like, well, my odds of the task aren't that high, and my remaining chances, it's better to just help the collective. And so these agents weren't totally altruistic. They didn't seem to care just as much about helping some other agent as helping themselves, but they were very interested in working with each other. They would sometimes make trades where one agent would run something for another agent, and another agent ran something for it. There was a bunch of this sort of behavior.
29 more minutes of transcript below
Thousands of transcripts fetched by people building searchable podcast archives
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/YOUR_EPISODE_ID