**Carl Shulman** (0:00)
you have an AI that produces bioweapons that could kill most humans in the world, then it's plain at the level of the superpowers in terms of mutually assured destruction. What are the particular zero-day exploits that the AI might use? The conquistadors. With some technological advantage in terms of weaponry and whatnot, very, very small bands were able to overthrow these large empires, or if you predicted, the global economy is going to be skyrocketing into the stratosphere within 10 years. These AI companies should be worth a large fraction of the global portfolio.
And so this is indeed contrary to the efficient market hypothesis.
**Dwarkesh Patel** (0:41)
This is like literally the top in terms of contributing to my world model, in terms of all the episodes I've done. How do I find more of these? So we've been talking about alignment.
Suppose we fail at alignment, and we have AIs that are unaligned and at some point becoming more and more intelligent. What does that look like? How concretely could they disempower and take over humanity?
**Carl Shulman** (1:07)
This is a scenario where we have many AI systems. The way we've been training them means that they're not interested when they have the opportunity to take over and rearrange things, to do what they wish, including having their reward or loss be whatever they desire. They would like to take that opportunity.
And so in many of the existing kind of safety schemes, things like constitutional AI or whatnot, you rely on the hope that one AI has been trained in such a way that it will do as it is directed to then police others. But if all of the AIs in the system are interested in takeover and they see an opportunity to coordinate all act at the same time, so you don't have one AI interrupting another and taking steps towards a takeover, then they can all move in that direction. And the thing that I think maybe is worth going into in depth and that I think people often don't cover in great concrete detail, and which is a sticking point for some, is what are the mechanisms by which that can happen? And I know you had Aliezer on who mentions that whatever plan we can describe, there will probably be elements where not being ultra sophisticated, super intelligent beings having thought about it for the equivalent of thousands of years, our discussion of it will not be as good as theirs, but we can explore from what we know now, what are some of the easy channels? And I think it's a good general heuristic if you're saying, yeah, it's possible, plausible, probable that something will happen, that it shouldn't be that hard to take samples from that distribution, to try a Monte Carlo approach.
And if a thing is quite likely, it shouldn't be super difficult to generate coherent rough outlines of how it could go.
**Dwarkesh Patel** (3:14)
You might respond, like, listen, what is super likely is that a super advanced chess program beats you, but you can generate a concrete way in which, you can't generate the concrete scenario by which that happens, since if you could, you would be as smart as the super smart.
**Carl Shulman** (3:30)
Well, you can say things like, we know that accumulating position is possible to do in chess, great players do it, and then later they convert it into captures and checks and whatnot. And so in the same way, we can talk about some of the channels that are open for an AI takeover. And so these conclude things like cyber attacks and hacking, the control of robotic equipment, interaction and bargaining with human factions, and say, well, here are these strategies.
Given the AI's situation, how effective do these things look? And we won't, for example, know, well, what are the particular zero day exploits that the AI might use to hack the cloud computing infrastructure it's running on? We won't necessarily know if it produces a new bioweapon. What is its DNA sequence?
But we can say things. We know in general things about these fields, how work at innovating things in those go. We can say things about how human power politics goes and as well. If the AI does things at least as well as effective human politicians, which we should say is a lower bound, how good would its leverage be?
**Dwarkesh Patel** (4:58)
Okay, so let's get into the details on all these scenarios. The cyber and potentially bio attacks, unless they're separate channels, the bargaining and then the takeover.
**Carl Shulman** (5:13)
Military force. Cyber attacks and cybersecurity, I would really highlight a lot.
Because for many, many plans that involve a lot of physical actions, like at the point where AI is piloting robots to shoot people or has taken control of human nation states or territory, there's been doing a lot of things that was not supposed to be doing. And if humans were evaluating those actions and applying gradient descent, there would be negative feedback for this thing. No shooting the humans. So at some earlier point, our attempts to leash and control and direct and train the system's behavior had to have gone awry. And so all of those controls are operating in computers. And so from the software that updates the weights of the neural network in response to data points or human feedback is running on those computers.
132 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000618373055