Topics: Society & Culture, Health & Fitness
**Chris Williamson** (0:00)
Hello, friends, welcome back. My guest today is Brian Christian, and we're talking about AI's scary challenge, The Alignment Problem. Let's say that you have a computer system and you want it to do X. You give it a set of examples and you say, well, do that. What could go wrong? Well, lots, apparently. And the implications are quite terrifying. So today, expect to learn why it's so hard to code an artificial intelligence to do what we actually want it to, how a robot cheated at the game of football, how human biases can be absorbed by AI systems, the most effective way to teach machines to learn, the danger if we don't get the alignment problem fixed, and much more. This is a topic I think everybody should be educated on. As the world becomes increasingly dependent on artificial intelligence, you need to understand the implications if we have a civilization built upon machine learning when we don't know that the machines are actually going to be aligned with what we want them to do. So today, hopefully, your curiosity will be satisfied and your pants might be pooped a little bit because it's quite, quite grave and quite scary if we don't get this right. But now it's time to learn about the alignment problem with Brian Christian.
What does the quote, premature optimization is the root of all evil mean?
**Brian Christian** (1:46)
So this line comes from Donald Knuth, who is one of the, I think of him as kind of like the Yoda of computer science, just dispensing these gems of wisdom. And there are many, I think many, like many aphorisms, you can take it in a number of different directions. One of the ways that I think about it is a lot of the way that we make progress in math and computer science is through models. You make a model that sort of approximates the phenomenon that you're trying to deal with. There's a great quote from Peter Norvig, another one of these luminaries in computer science. He's quoting someone from NASA saying, our job was not to land on Mars, it was to land on the mathematical model of Mars provided to us by the geologists.
This idea that premature optimization is the root of all evil, I think if you mistake the map for the territory, so to speak, if you forget that there's a gap between your model and what the reality actually is, then you can commit yourself to a set of assumptions that are later going to bite you. This is the sort of thing that people who are worried about AI safety, this is what keeps them up at night.
**Chris Williamson** (3:13)
What is the alignment problem? That's what we're going to be talking about today. We might as well define our terms.
**Brian Christian** (3:19)
Yeah. The alignment problem is this idea in AI and machine learning of the potential gap between what your intention is when you build an AI system or a machine learning system and the actual objective that the system has. So it's the potential misalignment, so to speak, between your intention, your expectation, how you want the system to behave and what that system ultimately ends up doing.
**Chris Williamson** (3:49)
Why does it matter?
**Brian Christian** (3:53)
I mean, this is a fear that has existed in computer science going back to at least 1960 So Norbert Wiener, the MIT cyberneticist, was writing about this, and he says, if we use to achieve some purpose a mechanical agency that we can't interfere with once we've started it, then we had better be quite sure that the purpose that we put into the machine is the thing that we really want. And I think a lot of people increasingly since like 2014, it has become more and more mainstream within the computer science community itself to think of this as one of the most significant challenges facing the field as we sort of enter this era of AI, that we may develop systems which, you know, we with the best of intentions try to encode some objective into the system. The system with, you know, sort of the best of intentions attempts to do what it thinks we want. But there's some fundamental misalignment and that results in, you know, whatever the harm may be, whether it's, you know, dark skinned people not getting recognized by, you know, facial recognition system or disparities in the way that parole is being dealt with, you know, at a societal level. It could be self-driving cars that failed to recognize jaywalkers, and so they kill anyone who's crossing in the middle of the street because there were no jaywalkers in their training data. All the way through to some of the so-called existential risks, the idea that we may actually throw society as a whole off the rails by some system with, you know, enough power to shape the course of human civilization, but without the appropriate wisdom to know exactly what to be doing.
55 more minutes of transcript below
Thousands of transcripts fetched by people building searchable podcast archives
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/YOUR_EPISODE_ID