Paul Christiano — Preventing an AI takeover artwork

Paul Christiano — Preventing an AI takeover

Dwarkesh Podcast

October 31, 2023

Paul Christiano is the world’s leading AI safety researcher. My full episode with him is out! We discuss: - Does he regret inventing RLHF, and is alignment necessarily dual-use?
Speakers: Dwarkesh Patel, Paul Christiano
**Dwarkesh Patel** (0:00)
, today, I have the pleasure of interviewing Paul Christiano, who is the leading AI safety researcher. He's the person that labs and governments turn to when they want feedback and advice on their safety plans. He previously led the language model alignment team at OpenAI, where he led the invention of RLHF, and now he is the head of the alignment research center, and they've been working with the big labs to identify when these models will be too unsafe to keep scaling.
Paul, welcome to the podcast.

**Paul Christiano** (0:33)
Thanks for having me. Looking forward to talking.

**Dwarkesh Patel** (0:35)
Okay, so first question, and this is a question I've asked Holden, Ilya, Dario, and none of them are going to be a satisfying answer.
Give me a concrete sense of what a post-AGI world that would be good would look like. Like how are humans interfacing with the AI? What is the economic and political structure?

**Paul Christiano** (0:52)
Yeah, I guess this is a tough question for a bunch of reasons. Maybe the biggest one is concrete. And I think it's just, if we're talking about really long spans of time, then a lot will change. And it's really hard for someone to talk concretely about what that will look like without saying really silly things. But I can mention some guesses or fill in some parts.
I think this is also a question of how good is good. Like often I'm thinking about worlds that seem like kind of the best achievable outcome or a likely achievable outcome.
So I am very often imagining my typical future has sort of continuing economic and military competition amongst groups of humans. I think that competition is increasingly mediated by AI systems. So for example, if you imagine humans making money, it'll be less and less worthwhile for humans to spend any of their time trying to make money or any of their time trying to fight wars. So increasingly the worlds you imagine is one where AI systems are doing those activities on behalf of humans. So like I just invest in some index fund and a bunch of AI's are running companies and those companies are competing with each other, but that is kind of a sphere where humans are not really engaging much.
The reason I gave this like how good is good caveat is like it's not clear if this is the world you'd most love. Like I'm like, yeah, the world, I'm leading with like the world still has a lot of war and it's a lot of economic competition and so on. But maybe what I'm trying to, what I'm most often thinking about is like, how can a world be reasonably good like during a long period where those things still exist? I think like in the very long run, I kind of expect something more like strong world government rather than just this like status quo. But that's like a very long run. I think there's like a long time left of like having a bunch of states and a bunch of different economic powers.

**Dwarkesh Patel** (2:27)
One word government, why do you think that's the transition that's likely to happen at some point?

**Paul Christiano** (2:32)
Yeah, so again, at some point I'm imagining or I'm thinking of like the very broad sweep of history. I think there are like a lot of losses, like wars are very costly thing. We would all like to have fewer wars. If you just ask like, what is humanity's long-term future like?
I do expect to drive down the rate of war to very, very low levels eventually. It's sort of like this kind of technological or social technological problem of like, sort of how do you organize society? How do you navigate conflicts in a way that doesn't have those kinds of losses?
And in the long run, I do expect us to succeed. I expect it to take kind of a long time subjectively. I think an important fact about AI is just like doing a lot of cognitive work and more quickly getting you to that world, more quickly or figuring out how do we set things up that way?

**Dwarkesh Patel** (3:09)
Yeah, the way Carl Shulman put it on the podcast is that you would have basically a thousand years of intellectual progress or social progress in a span of a month or whatever when the intelligence explosion happens. More broadly, so the situation where we have these AIs for managing our hedge funds and managing our factories and so on, that seems like something that makes sense when the AI is human level.
But when we have superhuman AIs, do we want gods who are enslaved forever? In a hundred years, what is the situation we want?

212 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000633226398