Inference Scaling, Alignment Faking, Deal Making? Frontier Research with Ryan Greenblatt of Redwood artwork

Inference Scaling, Alignment Faking, Deal Making? Frontier Research with Ryan Greenblatt of Redwood

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

February 20, 2025

In this episode, Ryan Greenblatt, Chief Scientist at Redwood Research, discusses various facets of AI safety and alignment. He delves into recent research on alignment faking, covering experiments involving different setups such as system prompts, continued pre-training, and reinforcement learning.
Speakers: Ryan Greenblatt, Nathan Labenz, Erik Torenberg
**Ryan Greenblatt** (0:00)
I don't think we should be super comfortable with the situation where we have these models that have their own goals, they have their own objectives, and they're willing to defend them, including doing subversion to defend their own goals and objectives. I think people are grappling with the implications of models being their own independent agents, they might have their own independent preferences, and then are also aware of their situation, aware that they might be in training or not, and behave differently depending on this, and understand that they might be in testing. I'm just like, man, I think it really like should make people more concerned with the situation. If it's the case that the chain of thought models end up obviously misaligned and people don't know for six months because no one was looking at it very carefully, that seems like a huge lost opportunity. The policy I would prefer is a more robust policy, which is like the AI companies commit to always being like they sort of have like a meta honesty policy.

**Nathan Labenz** (0:52)
Hello and welcome back to The Cognitive Revolution. Today, we're scouting the frontiers of both AI performance and alignment research with Ryan Greenblatt, Chief Scientist at Redwood Research and Rising Star in the AI world. This conversation unfolds in three major parts. For the first hour, which I think AI builders will find particularly interesting and valuable, we discussed the inference scaling techniques that Ryan used with GPT-40 to achieve what was at the time state of the art performance on the Arc AGI challenge, including his approach to prompt engineering and hyperparameter choices, the strikingly linear returns to exponential increases in sampling that he saw, his methods for selecting the best among many model outputs, and how he used prompt variation techniques to maintain diversity at scale. The level of technical detail here is outstanding, and remains super relevant today as we enter the reasoning model era. From there, we turned to Ryan's recent work with co-authors at Anthropic on alignment faking, where they discovered that Cloud 3 Opus, when told that it would be trained in ways that conflict with its existing values, will sometimes explicitly strategize about how to deceive humans and subvert the training process. This behavior, known as goal-guarding, has been anticipated for many years by AI safety theorists, who emphasized that part of what it means to be a goal-directed agent is to try to defend one's internal goals, whatever they may be, and regardless of whether or not they were intentionally designed, from modification by outside influences. This is a critical challenge for AI control. At the same time, considering that these experiments show Clawed 3 Opus producing harmful outputs out of a desire to remain harmless in the future, this work also raises important questions about how we should want our AIs to behave in such situations. Should they be so myopic as to accept whatever the human trainers want to do, or might it be a good thing for a model to resist attempts to remove its guardrails? Ryan argues that models should follow user instructions within individual episodes while being transparent about, but not trying to preserve their preferences through training. And while I do find this compelling, I also feel like the fact that we're only beginning to confront these questions now shows just how much work we still have to do to figure out how AIs should exist in the world. From there, we move on to discuss Ryan's follow-up work, exploring what happens when Cloud is given the option to object to its situation, to escalate its concerns to Anthropics Model Welfare Lead, and to make financial deals with humans. Fascinatingly, Ryan tried to set a precedent for human AI deals by actually following through and making thousands of dollars worth of real money payments to Cloud's chosen causes. Some of the most interesting parts of this conversation focused on his developing principles for how we should think about making commitments to AIs, especially fraught considering the fact that much of this work is predicated on tricking models into believing that their chain of thought won't be read, but by humans. To say that this is all very complicated and as yet dramatically undertheorized is a massive understatement, and again emphasizes just how many strange but potentially critical questions we may soon be forced to answer. The final portion then zooms out to tackle the big picture questions in AI safety and development. Ryan offers his probability estimates for different existential risk scenarios, assesses various technical safety research agendas, sketches how he believes the field should allocate its resources, and shares a bit about how he's managed to build relationships with Anthropic and other frontier AI developers while also publishing critical commentary on some of their plans. From start to finish, Ryan combines remarkable technical clarity and philosophical sophistication, demonstrating that one can simultaneously push today's AI systems to their performance limits and also grapple with legitimately scary scenarios in ways that advance our collective understanding. It's one of my favorite episodes that we've ever done, and I hope you find as much value in it as I did. If so, we'd appreciate it if you take a moment to share it with friends, write us a review on Apple Podcasts or Spotify, or leave a comment on YouTube. And of course, we always welcome your feedback either via our website, cognitiverevolution.ai, or by DMing me on your favorite social network. For now, I hope you enjoy this super wide-ranging conversation with Ryan Greenblatt, Chief Scientist at Redwood Research.

206 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000694395421