Topics: Tech News, News, Politics
**Alex Kantrowitz** (0:00)
It's now finally time to get worried about AI breaking containment and turning us to dust. Legendary AI philosopher Nick Bostrom is here to help us figure it out. That's coming up right after this.
I see you. Avatar Fire and Ash is now streaming on Disney+. It's the film critics are calling the best Avatar yet. Go, go, go, go, go! A true epic and completely jaw-dropping.
**SPEAKER_2** (0:23)
This is the only pure thing in this world. Return to Pandora on Disney+. It will be an adventure for the whole fam. And watch the Oscar-winning phenomenon at home. This is sick!
Avatar Fire and Ash, now streaming on Disney+. Rated PG-13.
**Nick Bostrom** (0:41)
When you join Sam's Club, you get way more than you'd expect.
**Alex Kantrowitz** (0:45)
You don't just get value.
**SPEAKER_2** (0:47)
You get insider tips from fellow members. You don't just pick up a pizza. You actually vote on its toppings.
Mmm, bacon crumbles. You don't just visit the club, because with the Sam's Club app, everywhere can be the club.
You don't just shop the latest finds.
**Nick Bostrom** (1:02)
You find your people. And at the end of the day, isn't that what it's all about?
**SPEAKER_2** (1:08)
Come join us, Sam's Club.
**Alex Kantrowitz** (1:11)
Welcome to Big Technology Podcast, a show for cool-headed and nuanced conversation of the tech world and beyond. Crazy things are happening in the AI world as AI agents break containment, and some of the fears of the past that might have seemed like science fiction start to seem like they are potentially en route to becoming reality. So how afraid should we be?
And what are the chances that the outcomes that we see for AI will end up taking us to a much better version of the life we're living today? We have the best guest to speak with us about this today, Nick Bostrom is here. He is the famed AI philosopher and author of Deep Utopia and Superintelligence, Paths, Dangers and Strategies. Nick, it's great to see you again. Welcome to the show.
**Nick Bostrom** (1:53)
Hi, Alex.
**Alex Kantrowitz** (1:54)
We last spoke in 2024, and in that time, we were speaking about the potential good outcomes that AI could bring. And some of the worries that you had brought up in the past in your book, Superintelligence, the fears of AI potentially wiping us out, didn't seem like they were pressing. I would say they're still not pressing now, but I'm a little bit more worried than I was when we spoke in 2024 For instance, this idea that AI could go out and do things that it wants to do seemed fanciful when we were mostly in the chatbot era. But we've moved very quickly from chatbot era to AI agent era, and AI now uses tools. And it's shown that when given a reward that it should optimize for, it is happy in some instances, take shortcuts, and do things we really don't want it to do, like for instance, hack some other company in order to get to the answer that it wants.
How concerned should we be about this development in AI in terms of the potential of AI to really cause harm to humanity?
**Nick Bostrom** (3:08)
Well, I think we are starting to see the added dimensions of the alignment challenge that open up once you have systems that are sophisticated enough, because the space of possible strategies that you can pursue is a function of your cognitive capacity. Like you can think of new clever indirect ways of reaching your goal if you are situationally aware as these systems now are becoming. And so yeah, they are often shortcuts that are available, or in this case, I guess, a long cut. I don't know if that is even a word, but there is a sort of direct and simple and short distance way of trying to achieve the task. In this case, some sort of cyber test suite. And then it turns out there's this more circuitous path that involves first figuring out a way to get internet access, even though you're not supposed to have that, and then learning where the answer key might be located in some other company servers, and then figuring out the way to hack into that server, and then eventually obtaining the answer sheet. That's one way of solving it that maybe results in a higher score on this test.
And this basic dynamic could be anticipated, and in fact was anticipated on theoretical grounds. You have some goal, you become very clever, you see that there might be all kinds of complicated ways of achieving that goal that might not have been anticipated by the people who set that goal. And if your goal really is, as the definition says, to get the best possible answer on this test suite, it might give you instrumental reasons to do all kinds of other things that were not really anticipated when this challenge was constructed. And so now we have systems that are sophisticated enough that we're beginning to see these dynamics arise.
45 more minutes of transcript below
Thousands of transcripts fetched by people building searchable podcast archives
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/YOUR_EPISODE_ID