Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI artwork

Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI

Latent Space: The AI Engineer Podcast

June 19, 2025

Solving Poker and Diplomacy, Debating RL+Reasoning with Ilya, what’s *wrong* with the System 1/2 analogy, and where Test-Time Compute hits a wall Full Video Episode Timestamps 00:00 Intro – Diplomacy, Cicero & World Championship 02:00 Reverse Centaur: How AI Improved Noam’s Human Play 05:00 Turing...
Speakers: Alessio, Spooks, Noam Brown
**Alessio** (0:05)
Hey, everyone. Welcome to the Latent Space Podcast. This is Alessio, partner and CTO at Decibel, and I'm joined by my co-host, Spooks, founder of SmallAI.

**Spooks** (0:12)
Hello, hello. And we're here recording on a holiday Monday with Noam Brown for an OpenAI. Welcome. Thank you. So glad to have you finally join us. A lot of people have heard you. You've been rather generous of your time on the podcast, Lex Friedman, and you've done a TED talk recently just talking about the thinking paradigm.
But I think maybe perhaps your most interesting recent achievement is winning the World Diplomacy Championship. In 2022, you built Cicero, which was top 10% of human players. I guess my opening question is, how has your diplomacy playing changed since working on Cicero and now personally playing it?

**Noam Brown** (0:53)
When you work on these games, you kind of have to understand the game well enough to be able to debug your bot. Because if the bot does something that's really radical and that humans typically wouldn't do, you're not sure if that's a mistake, or if it's a bug in the system, or it's actually just the bot being brilliant. When we were working on Diplomacy, I did this deep dive trying to understand the game better. I played in tournaments, I watched a lot of tutorial videos and commentary videos on games, and over that process, I got better. Then also seeing the bot, the way it would behave in these games, sometimes it would do things that humans typically wouldn't do. And that taught me about the game as well. When we released Cicero, we announced it in late 2022 I still found the game really fascinating, and so I kept up with it, I continued to play, and that led to me winning the championship, the world championship in 2025, so just a couple months ago.

**Spooks** (1:46)
There's always a question of centaur systems where humans and machines work together. Was there an equivalent of what happened in Go where you updated your play style?

**Noam Brown** (1:55)
If you're asking if I used Cicero when I played in the tournament, the answer is no. Seeing the way the bot played and taking inspiration from that, I think did help me in the tournament.

**Spooks** (2:05)
Yeah. Do people now ask Turing questions every single time when they're playing Diplomacy?

**Noam Brown** (2:11)
Ask to try to tell if the person they're playing with is a bot or a human.

**Spooks** (2:17)
Yeah, that's the one thing you were worried about when you started.

**Noam Brown** (2:20)
It was really interesting when we were working on Cicero because we didn't have the best language models. We were really bottle-necked on the quality of the language models. And sometimes the bot would do, would say, bizarre things.
99% of the time it was fine, but then every once in a while, it would say this really bizarre thing. It would just hallucinate about something. Somebody would reference something that they said earlier in a conversation with the bot, and the bot would be like, I have no idea what you're talking about. I never said that. And then the person would be like, look, you could just scroll up in the chat and it's literally right there. And the bot would be like, no, you're lying.

**Spooks** (2:49)
It's Windows.

**Noam Brown** (2:51)
And when it does these kinds of things, people just shrug it off as like, oh, that's just the person's tired, or they're drunk, or whatever, or they're just trolling me. But I think that's because people weren't looking for a bot. They weren't expecting a bot to be in the games. We were actually really scared because we were afraid that people would figure out at one point that there's a bot in these games, and then they would just always be on the lookout for it. And if you're looking for it, you're able to spot it. That's the thing. So I think now that it's announced and that people know to look for it, I think they would have an easier time spotting it. Now, that said, the language models have also gotten a lot better since 2022

**Spooks** (3:27)
It's adversarial. Yeah.

**Noam Brown** (3:28)
So at this point, the truth is, GPD 4 and O3, these models are passing the Turing test. So I don't think they can really ask that many Turing complete questions that would actually make a difference.

**Alessio** (3:39)
And Cicero was very small, like 27B, right?

**Noam Brown** (3:42)
It was a very small language model, yeah. It was one of the things we realized over the course of the project that, like, oh, yeah, you really benefit a lot from just having larger language models.

80 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000748427701