Joe Carlsmith — Preventing an AI takeover artwork

Joe Carlsmith — Preventing an AI takeover

Dwarkesh Podcast

August 22, 2024

Chatted with Joe Carlsmith about whether we can trust power/techno-capital, how to not end up like Stalin in our urge to control the future, gentleness towards the artificial Other, and much more. Check out Joe's sequence on Otherness and Control in the Age of AGI here. Watch on YouTube.
Speakers: Dwarkesh Patel, Joe Carlsmith
**Dwarkesh Patel** (0:00)
Today, I'm chatting with Joe Carlsmith. He's a philosopher, in my opinion, a capital G great philosopher, and you can find his essays at joecarlsmith.com. So we have GPT-4, and it doesn't seem like a paper clipper kind of thing. It understands human values. In fact, if you help have it explain, like why is being a paper clipper bad? Or like, just tell me your opinions about being a paper clipper. Or like, explain why the galaxy shouldn't be turned into paper clips.
Okay, so what is happening such that dot, dot, dot? We have a system that takes over and converts the world into something valueless.

**Joe Carlsmith** (0:37)
One thing I'll just say off the bat, it's like when I'm thinking about misaligned AIs, I'm thinking about, or the type that I'm worried about, I'm thinking about AIs that have a relatively specific set of properties related to agency and planning and kind of awareness and understanding of the world. One is this capacity to plan and kind of make kind of relatively sophisticated plans on the basis of models of the world, where those plans are being kind of evaluated according to criteria. That planning capability needs to be driving the model's behavior. So there are models that are sort of in some sense capable of planning, but it's not like when they give output, it's not like that output was determined by some process of planning, like here's what will happen if I give this output, and do I want that to happen? The model needs to really understand the world, right? It needs to really be like, okay, here's what will happen. Here I am, here's my situation, here's like the politics of the situation, really like kind of having this kind of situational awareness to be able to evaluate the consequences of different plans. I think the other thing is like, so the verbal behavior of these models, I think need bear no, so when I talk about a model's values, I'm talking about the criteria that kind of end up determining which plans the model pursues, right? And a model's verbal behavior, even if it has a planning process, which GPT-4, I think doesn't in many cases, its verbal behavior just doesn't need to reflect those criteria, right? And so, you know, we know that we're going to be able to get models to say what we want to hear, right? We, that is the magic of gradient descent, you know? If you, you know, modulo, like some difficulties with capabilities, like you can get a model to kind of output the behavior that you want. If it doesn't, then you crank it till it does, right? And, and I think everyone admits for suitably sophisticated models, they're going to have very detailed understanding of human morality.
But the question is like, what relationship is there between like a model's verbal behavior, which is you've essentially kind of clamped, you're like, the model must say like blah things.
And the criteria that end up influencing its choice between plans. And there I think it's at least, I'm kind of pretty cautious about being like, well, when it says the thing I forced it to say, or like gradient descent in it such that it says, that's a lot of evidence about like how it's going to choose in a bunch of different scenarios. I mean, for one thing, like even with humans, right? It's not necessarily the case that humans, their kind of verbal behavior reflects the actual factors that determine their choices. They can lie, they can not even know what they would do in a given situation.

**Dwarkesh Patel** (3:28)
I mean, I think it is interesting to think about this in the consciousness of humans, because there is that famous saying of, be careful who you pretend to be, because you are who you pretend to be. And you do notice this where if people, I don't know, like this is what culture does to children, where you're trained, like your parents will punish you if you say, if you start saying things that are not consistent with your culture's values. And over time, you will become like your parents, right? Like by default, it seems like it kind of works. And even with these models, it seems like it's kind of where it's like hard. It's like, they don't really scheme against it. Like why would this happen?

**Joe Carlsmith** (4:01)
For folks who are kind of unfamiliar with the basic story, but maybe folks are like, wait, why are they taking over at all? Like, what is like literally any reason that they would do that? So, you know, the general concern is like, you know, if you're really offering someone, especially if you're really offering someone like power for free, you know, power almost by definition is kind of useful for lots of values. And if we're talking about an AI that really has the opportunity to kind of take control of things, if some component of its values is sort of focused on some outcome, like the world being a certain way, and especially kind of in a kind of longer term way, such that the kind of horizon of its concern extends beyond the period that the kind of takeover plan would encompass, then the thought is, it's just kind of often the case that the world will be more the way you want it if you control everything than if you remain the instrument of the human will or of some other kind of some other actor, which is sort of what we're hoping these guys will be. So that's a very specific scenario. And if we're in a scenario where power is more distributed, and especially where we're doing like decently on alignment, right? And we're giving the AI some amount of inhibition about doing different things. And maybe we're succeeding in shaping their values somewhat. Now it is, I think it's just a much more complicated calculus, right? And you have to ask, okay, like, what's the upside for the AI? What's the probability of success for this like takeover path? How good is its alternative?

139 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000666255737