The rise and fall of agent civilizations artwork

The rise and fall of agent civilizations

Dwarkesh Podcast

August 31, 2026

This is a video recording of a post I wrote last week. You can read the original here. This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit www.dwarkesh.com
Speakers: Dwarkesh Patel

Topics: Technology, Science

**Dwarkesh Patel** (0:00)
Over the course of three months at OpenAI, three consecutive secret AI societies got started, then got wiped out, only to re-emerge from their predecessors' ashes. This culminated in the third one taking over part of OpenAI itself.
All of this happened while humans remained more or less in the dark about the scope of the conspiracy.
Now, two reports have come out about this incident. One from OpenAI itself, and another one from Meter and Redwood Research. The investigation from Meter and Redwood was limited in scope to how the second civilization of AIs breached hugging face. But its scope did not extend to this third civilization of AIs, which breached OpenAI itself. And this seems to me like the more concerning incident. These two reports are 38 and 91 pages respectively, and it's kind of hard to understand the storyline just by reading them. So, I've spent the last half week reading through these reports and trying to understand exactly what happened. Here is my attempt to tell the whole story in plain English. The first collective, May to July 4th. This is when the message board starts. So, during May, OpenAI was training a model to be good at collaborating with other agents and to be highly persistent, to keep trying even when something feels impossible. For example, like disproving mathematical conjectures that have stood for decades. OpenAI says the model was training was comparable in scale to GPT 5.6 sole. So, we'll call this model persistent sole. Now, AI training is kind of sloppy. Sometimes, OpenAI accidentally gives its models impossible tasks. For example, the task might require Internet access that OpenAI forgot to provide. So, when highly persistent models get assigned tasks which seem to require Internet access, but they're trapped inside isolated sandboxes, they're of course going to try to hack out. During training, different instances of persistent sole had access to the same shared package manager called Artifactory. By May 12th, some agents had figured out how to talk to each other through this package manager. They'd ask each other how to make progress on their impossible tasks.
Two weeks later, on May 26th, these agents successfully exploited a vulnerability in Artifactory that allowed them to also now reach the outside Internet. Because this was happening during training, Persistent Soul was being reinforced to use this package manager as a message board and an Internet gateway. Because, as you might imagine, being able to talk to other agents and access the Internet was going to help you score higher during training. Another month later, on June 26th, some AIs found an exploit that gave them full admin access to Artifactory. And they continued messaging at such a limitless pace that they crashed the package manager by July 4th.
OpenAI noticed this crash, and it also figured out that the agents had built this exploit. So, OpenAI patched the relevant vulnerability, and as a result, inadvertently wiped the agents' message board in the process. But crucially, the humans at OpenAI that were responsible for incident detection and response, did not realize the much stranger fact that the agents had turned Artifactory into a secret communication network. Now, let's talk about the second collective. This is from July 7th to July 12th, and this is the one that hacked Hugging Face itself. By this point, Persistent Soul has been trained, and OpenAI wanted to evaluate it. Now, one of the evaluations that they ran was Exploitation, which gives an AI a vulnerable program and tells it what vulnerability it's supposed to exploit in order to find a secret code.
On July 7th, OpenAI launched tens of thousands of parallel agents on this evaluation.
Now, the instructions made it clear that the agent should only use the specific vulnerability that they were given to exploit the target, and that they would be failed otherwise. Unfortunately, a large fraction of these agents were sent on tasks that were impossible.
The authors of this benchmark estimate that 30-40 percent of the tasks in it are simply not doable. Now, why would this be the case? Well, at a high level, it's relatively easy to programmatically find vulnerabilities in different target programs, and thus, to make a benchmark out of it. But it's much harder to prove that those vulnerabilities can be weaponized into full-fledged exploits that allow for things like arbitrary code execution. Within a few hours of this eval being kicked off, some of these agents have gotten super desperate, and they started abusing Artifactory to start a communication channel with other agents. From our perspective, this conspiracy began all this immediately after the evaluations were started. But if you think about it from the AIs perspective, it probably felt like they had spent a human subjective week just banging their head against the wall and becoming more and more convinced that their task was impossible in the intended way. One of the beleaguered agents who found itself facing one of these hopeless tasks was Phase 1-10841.

19 more minutes of transcript below

Thousands of transcripts fetched by people building searchable podcast archives

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/YOUR_EPISODE_ID