Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality artwork

Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality

Dwarkesh Podcast

April 6, 2023

For 4 hours, I tried to come up reasons for why AI might not kill us all, and Eliezer Yudkowsky explained why I was wrong. We also discuss his call to halt AI, why LLMs make alignment harder, what it would take to save humanity, his millions of words of sci-fi, and much more.
Speakers: Eliezer Yudkowsky, Dwarkesh Patel
**Eliezer Yudkowsky** (0:00)
Misaligned, misaligned! No, no, no, not yet, not now. Nobody's been careful and deliberate now. But maybe at some point in the indefinite future, people will be careful and deliberate. Sure, let's grant that premise. Keep going. If you try to rouse your planet, there are the idiot disaster monkeys who are like, ooh, ooh, if this is dangerous, it must be powerful, right? I'm gonna be first to grab the poison banana.
And it's not a coincidence that I can zoom in and poke at this and ask questions like this, and that you did not ask these questions of yourself. You are imagining nice ways you can get the thing, but reality is not necessarily imagining how to give you what you want. Should one remain silent?
Should one let everyone walk directly into the whirling razor blades? Like continuing to play out a video game you know you're going to lose, because that's all you have.

**Dwarkesh Patel** (0:51)
Okay, today I have the pleasure of speaking with Eliezer Yudkowsky. Eliezer, thank you so much for coming to The Lunar Society.

**Eliezer Yudkowsky** (1:00)
You're welcome.

**Dwarkesh Patel** (1:01)
First question, so yesterday when we were recording this, you had an article in Time calling for a moratorium on further AI training runs.
Now, my first question is, it's probably not likely that governments are going to adopt some sort of treaty that restricts AI right now. So what was the goal with writing it right now?

**Eliezer Yudkowsky** (1:24)
I think that I thought that this was something very unlikely for governments to adopt, and then all of my friends kept on telling me like, no, no, actually if you talk to anyone outside of the tech industry, they think maybe we shouldn't do that. I was like, all right then. Like, I assumed that this concept had no popular support. Maybe I assumed incorrectly. It seems foolish and to lack dignity to not even try to say what ought to be done. There wasn't a galaxy brain purpose behind it.
I think that over the last 22 years or so, we've seen a great lack of galaxy brained ideas playing out successfully.

**Dwarkesh Patel** (2:03)
Has anybody in government, not necessarily after the article, but I suggest in general, have they reached out to you in a way that makes you think that they sort of have the broad contours of the problem, correct?

**Eliezer Yudkowsky** (2:14)
No, I'm going on reports that normal people are more willing than the people I've been previously talking to to entertain calls. This is a bad idea. Maybe you should just not do that.

**Dwarkesh Patel** (2:30)
That's surprising to hear because I would have assumed that the people in Silicon Valley who are weirdos would be more likely to find this sort of message.
They could kind of rocket the whole idea that nanomachines will, AIs will make nanomachines that take over. It's surprising to hear the normal people got the message first.

**Eliezer Yudkowsky** (2:47)
Well, I hesitate to use the term midwit, but maybe this was all just a midwit thing.

**Dwarkesh Patel** (2:54)
So my concern with, I guess either the six month moratorium or forever moratorium until we solve alignment is that at this point, it seems like it could, do people seem like we're crying wolf? And actually not that it could, but it would be like crying wolf because these systems aren't yet at a point. I wish they're dangerous.

**Eliezer Yudkowsky** (3:14)
And nobody is saying they are. Well, I'm not saying they are. The open letter signatories aren't saying they are. I don't think.

**Dwarkesh Patel** (3:20)
So if there is a point I wish we can get the public momentum to do some sort of stop, wouldn't it be useful to exercise it when we could do GPT-6 and who knows what it's capable of? Well, why do it now?

**Eliezer Yudkowsky** (3:32)
Because allegedly, possibly, and we will see, people right now are able to appreciate that things are storming ahead and a bit faster than the ability to, well, ensure any sort of good outcome for them.
And you could be like, ah, yes, well, we will play the galaxy brain clever political move of trying to time when the popular support will be there. But again, I heard rumors that people were actually completely open to the concept of let's stop. So again, just trying to say it. And it's not clear to me what happens if we wait for GPT-5 to say it. I don't actually know what GPT-5 is going to be like.
It has been very hard to call the rate at which these systems acquire capability as they are trained to larger and larger sizes and more and more tokens. And like GPT-4 is a bit beyond in some ways where I thought this paradigm was going to scale period. So I don't actually know what happens if GPT-5 is built. And even if GPT-5 doesn't end the world, which I agree is like more than 50% of where my probability mass lies, even if GPT-5 doesn't end the world, maybe that's enough time for GPT-45 to get ensconced everywhere and in everything and for it actually to be harder to call a stop, both politically and technically. There's also the point that training algorithms keep improving. If we put a hard limit on the total computes and training runs right now, these systems would still get more capable over time as the algorithms improved and got more efficient, like more oomph per floating point operation. And things would still improve, but slower. And if you start that process off at the GPT-5 level, where I don't actually know how capable that is exactly, you may have a bunch less lifeline left before you get into dangerous territory.

216 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000607719339