Historic AI Developments & the Emerging Shape of Superintelligence, from the Consistently Candid Podcast artwork

Historic AI Developments & the Emerging Shape of Superintelligence, from the Consistently Candid Podcast

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

March 8, 2025

In this episode, we discuss the significant advancements and challenges in the field of artificial intelligence over the past year. From breakthroughs in reinforcement learning to unexpected behaviors in fine-tuned models, we cover a wide range of topics that are shaping the AI landscape.
Speakers: Nathan Labenz, Sarah Hastings Woodhouse
**Nathan Labenz** (0:00)
I sort of suspect that in the like zoomed out history, this might appear to be like a critical threshold. You know, everybody was kind of scaling these base models, then somebody figured out that you could also scale inference compute. And then it became clear that it's actually pretty easy to do that. And then what's that going to produce? You know, it seems to me that it's likely to produce a lot of kind of weird AIs, because reinforcement learning also famously gives rise to strange behavior. It just seems like governance gets a lot harder in this world where distributed training works and where the post-training that really shapes the AIs behavior and like practical utility and how they're going to show up in the world is actually become like quite cheap. A race to powerful AGI between the US and China is one of the worst situations I can imagine that could lead to catastrophic outcome. I would say math and coding in particular are like almost undoubtedly going to hit superhuman levels in the next probably 2025, certainly seems like by 2026
Hello, and welcome back to The Cognitive Revolution. Today, I'm pleased to share a cross post of my appearance on Consistently Candid with host Sarah Hastings Woodhouse. This was my second episode with Sarah. Nine months ago, she was just getting into AI and still making sense of the fundamentals. Today, as you'll hear, she's developed a strong sense for which AI stories really matter and also does an excellent job of summarizing notable research results. Together, we unpack a number of stories that I believe history will judge to be among the most important of the last nine months. We start with the recent revelation that reinforcement learning can be relatively easily applied to sufficiently powerful base language models and the reasoning capabilities that this has unlocked. We then move on to consider the rise of distributed training, which especially as combined with the inference-heavy nature of reinforcement learning makes it possible for all sorts of moderately resourced organizations and distributed groups to apply reinforcement learning to any objective they might like. We also discuss the shift in rhetoric among American AI leaders toward embracing an AI arms race with China and get into a couple important recent AI alignment results, including the alignment faking paper that we covered in-depth in our episode with Ryan Greenblatt and also the very viral emergent misalignment paper from Aline Evans group, to which I made a minor contribution and on which I was honored to be included as a co-author. Perhaps most interesting for regular listeners, for the first time publicly, I offer a sketch of the form that I expect early superintelligence to take, in the base case, over the next few years. The upside is that AI systems' ability to develop intuitive physics across many different problem spaces, like material science, protein folding, cell biology, and many, many more, especially as combined with reasoning abilities, suggests a pretty clear path to an exponentially growing number of eureka moments from AI systems, which really could accelerate science to the point that we achieve a century's worth of progress in just the next few years. At the same time, on the downside are still nascent understanding of how these systems work, the rate at which we continue to be surprised by their outputs, and the growing body of evidence suggesting that frontier models are increasingly willing to deceive and otherwise scheme against their human users to achieve their own goals and protect their own values. All suggest that we will see lots of instances of bad behavior and will need to invest heavily in control measures along the way. This vision of superintelligence and also the vision of drop-in AI knowledge workers that I sketch out toward the end are inherently more forward-looking and speculative than my usual material. And as such, I really want your feedback on this episode in particular. Thanks to your consistent engagement with the feed, I'm getting more and more invitations to speak to business and general audiences about AI. And this picture of superintelligence, which I hope makes the potential for frontier discovery tangible and the reality of bad behavior plain, is becoming a consistent framing for my talks. So, what do you think? Am I missing something? Am I overstating any of the achievements or the possible benefits? Or perhaps understating any of the demonstrated issues? I really want to make sure that I'm sharing the most accurate and up-to-date understanding that I possibly can. So please do let me know if you think I'm getting anything wrong. With that, I hope you enjoy this review of the most important AI stories of recent months and this possible preview of the future of AI transformation. From the Consistently Candid Podcast, with Sarah Hastings-Woodhouse.

97 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000698411019