roon's Heroic Duty: Will "the Good Guys" Build AGI First?  (from Doom Debates) artwork

roon's Heroic Duty: Will "the Good Guys" Build AGI First? (from Doom Debates)

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

December 28, 2024

In this episode of The Cognitive Revolution, Nathan shares a fascinating cross-post from Doom Debates featuring a conversation between Liron Shapira and roon, an influential Twitter Anon from OpenAI's technical staff.
Speakers: Nathan Labenz, Liron Shapira, roon
**SPEAKER_1** (0:02)
Before we dive into today's episode, I want to tell you about a new show from Turpentine called Modern Relationships. On the season ahead, I sit down with power couples in tech and leading relationship thinkers to explore how ambitious people actually make partnerships work. Whether you're dating in a relationship or just curious how technology is reshaping modern love, I think you'd enjoy this on your feed. Our first episode features founders funds Adelian Asparuhov and tech researcher Nadya Asparuhov, who take us through their evolution from dating to marriage to parenthood, with absolutely no filter on the challenges and growth along the way. You can find Modern Relationships wherever you get podcasts. Now, on to today's episode.

**Nathan Labenz** (0:37)
Hello, and welcome back to The Cognitive Revolution. Today, I'm excited to share a cross post from Doom Debates by Liron Shapira, featuring a discussion between Liron and Rune, a widely respected and highly influential Twitter anonymous account known to be powered by a member of OpenAI's technical staff. What makes this conversation particularly valuable in my opinion is the window it provides into how people at OpenAI, and to a lesser extent, the other leading labs, are thinking about the near-term big-picture future of AI. While Rune is careful to say that he's not speaking on behalf of OpenAI, and I appreciate that this is a candid conversation not representing any official policy, I think it's nevertheless quite telling. So, what should you be listening for? Well, for starters, you'll notice that Rune is consistently quick to acknowledge and willing to grapple with AI's transformative potential. When asked whether AI might one day outperform Elon Musk at running companies, or whether it could match Terence Tao at mathematics, Rune takes these questions not as high-level metaphors, but as concrete empirical questions that we could very plausibly answer in the affirmative within the decade. At one point, he seems to take the concept of a technological singularity pretty much for granted. Saying that AGI is coming very soon, while by contrast, highly capable humanoid robots will take longer, by which he means maybe just two to three more years. With that in mind, it makes sense that Rune agrees that it would be logically incoherent to believe AIs will become super powerful without also being willing to confront the reality of tail risks, to endorse the famous AIX risk statement, which was signed by OpenAI leadership and which Rune describes as a very low bar, and to predict that Meta will eventually stop open sourcing their models, noting that it becomes irresponsible at a certain level to release new models immediately. Yet at the same time, despite this clear-eyed view of AI's trajectory and likely future capabilities, Rune puts the probability of human extinction from AI causes at less than 1%.
This to me seems much less well supported by the current state of evidence, so why does he believe it? His optimism seems to rest on three key pillars. First, belief in quote-unquote alignment by default, through the learning of human priors, which we might also call human values, during the pre-training phase. Second, belief in the moderating effects of competition between AI systems. And third, and most notably, confidence that the good guys will develop powerful AI first. Now, to give alignment by default its due, I've said many times on this feed that our ability to create AI's that understand human values and act ethically, has dramatically exceeded my expectations, and now seems way more plausible than I would have thought just a couple of years ago. Of course, at the same time, I've also noted that it doesn't really happen by default. So-called purely helpful models without safety mitigations are, in my experience, and here I'm speaking specifically of GPT-4 early, often shockingly amoral. In any case, it was actually the last point that particularly caught my attention, because as you may recall from previous episodes, I am always very suspicious of any analysis that begins by identifying good guys and bad guys, and then proceeds to reason from there. History and all sorts of experimental psychology results show that people regularly fail to question such assumptions when others most need them to, sometimes with catastrophic effect. So when Ruhm discusses his sense of duty to use maximum technical and strategic skill in his work on AI, or says that it's a pretty cool thing to say that I eradicated polio, I think we get a glimpse into the main character mindset that I really worry is all too common among those building these transformative systems. Yes, they genuinely aspire to build AGI for the benefit of all humanity, and yes, they really are trying to do a good job and be responsible about it. But might they be contenting themselves a little too quickly with better us than them? And exactly what P-Doom are they asking the rest of humanity to accept anyway?

91 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000681931970