**Edward Hughes** (0:00)
What happens when you drop humans into a Claude 3.5 Society or a GPT-40 Society or some mix of society? Do the humans end up behaving differently? Where does the society end up? My expectation is that LLM Agents are going to become a big thing. Everyone thinks the 2025 is the Year of Agents. I agree.
**Aron Vallinder** (0:21)
The best way to create trust is to be in an environment where people are, in fact, trustworthy and sort of cooperate with you. And so I think we will have to have certain standards or regulations for how these interactions work that are sort of designed to create a trusting environment.
**Nathan Labenz** (0:43)
Hello, and welcome back to The Cognitive Revolution. Today, I'm excited to share my conversation with Edward Hughes, Researcher at Google DeepMind, and Aron Vallinder, an Independent Researcher and Pibbs Fellow, who recently published a fascinating paper exploring Cultural Evolution in Toy AI Societies, and studying which of today's popular large language models do and don't cooperate well enough to sustain positive-sum social norms over time. Using a classic behavioral economics experiment called the Donor Game, where agents choose how much of a valuable resource to donate to another agent, which in turn receives twice the amount that the first agent donated, they demonstrate striking differences in how leading language models develop and maintain cooperative norms across generations. The results? In a game in which a perfectly cooperative society could accumulate 32,000 units of the resource, Claude 3.5 Sonnet does by far the best, achieving 3 to 5,000 units and showing increasingly pro-social behavior over time. Whereas in comparison, Gemini 1.5 Flash cooperates only limitedly and achieves a few hundred units. And GPT-40 shows very minimal cooperation and almost no resource growth. Beyond the headline findings, we discussed the details of how they implemented cultural transmission between generations of AI agents, the crucial role of reputation, including how important it is that AI agents enforce cooperative norms by punishing and rewarding the punishment of defectors, and the results of early experiments mixing different models together in the same society. This work highlights important blind spots in our standard benchmark-centric approach to characterizing AI systems. And I hope it gets more people thinking about how social norms and cultural dynamics might quickly begin to change as we introduce large numbers of AI agents to human society. More broadly still, I hope it gets you asking critical questions about our AI future that nobody else has yet thought to ask. Importantly, this kind of research is uniquely accessible. Aron and Edward have open sourced their code to invite others to build on their work. And in general, especially now with AI coding assistance, this kind of research requires very little technical skill. If you're an economist or social scientist and you're inspired to explore this kind of work but need help getting started, please do not hesitate to reach out. I would be happy to help orient, connect, and advise you. As always, if you're finding value in the show, we'd appreciate it if you'd take a moment to share it with friends, write a review on Apple Podcasts or Spotify, or leave us a comment on YouTube. We welcome your feedback and suggestions too, either via our website, cognitiverevolution.ai, or by DMing me on your favorite social network. With that, I hope you enjoy this early glimpse of Cultural Evolution in AI Societies with Edward Hughes and Aron Vallinder. Aron Vallinder and Edward Hughes, authors of Cultural Evolution of Cooperation among LLM Agents, welcome to The Cognitive Revolution.
**Edward Hughes** (3:37)
Pleased to be here.
**Aron Vallinder** (3:39)
Thanks so much.
**Nathan Labenz** (3:40)
I'm really excited about this. You guys have put out some really interesting work. I think it's some of the earliest work in what I expect will be a fast growing and super interesting field of just asking the question, what happens when we have a lot of AIs running around? I have been honestly looking for more research in this domain because I feel like so many of us are in AI in general, right? Everything's happening so fast. So many people are focused on their individual project, their individual line of research, or even if they're just daily users, their sort of implicit model of the world is so often like, mostly the world is as it is and is like normal, but I'm getting a little bit more productive with AI here and there. And I think, especially, we're talking on the same day that OpenAI debuted their new operator web agent, I think we're actually headed for probably a lot more change than that when we get to the point when AIs are running around autonomously and there's a lot of them and they're starting to interact with each other and the world is going to adapt in all kinds of ways and boy are we not ready for that. So I really appreciate that you guys are starting to take some of the first bites out of that very big apple and want to take the time today to really dig in and make sure I understand the work that you've already done and get a little sense of where you're going and hopefully inspire other people to come join you because I think there's a lot there to be done. How does that sound?
77 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000691467188