How Researchers Test AI for Hidden Goals — Apollo Research artwork

How Researchers Test AI for Hidden Goals — Apollo Research

Machine Learning Street Talk (MLST)

July 31, 2026

Can an AI do the right thing for the wrong reason? Tim Scarfe speaks with Apollo Research’s Alexander Meinke, Axel Højmark and Jérémy Scheurer about Measuring Reward-Seeking via Contrastive Belief Updates, their new research with OpenAI.
Speakers: Jérémy Scheurer, Tim Scarfe, Alexander Meinke, Axel Højmark
**Jérémy Scheurer** (0:00)
So, they might realize, for instance, that they are being tested. They might start thinking about, ah, the greater wants me to do this. So, that's what I might then do.

**Tim Scarfe** (0:08)
So, the other day on our Discord server, Fable deleted 100 messages from Wendy. I'm very sorry about that, Wendy. Fable is just so over-eager, it's so adaptable, and it's weird because we thought this was what we wanted.

**Alexander Meinke** (0:19)
Okay, so we said that, behaviorally, a reward seeker looks totally the same as an aligned model. So, how do you tell the difference? The first thing that you might try is, can we just ask the model?

**Axel Højmark** (0:30)
And essentially, when you train these models to believe that graders reward task completion at all costs, they will, in a scenario where deception is needed to solve the task, they will break their promise 87% of the time.

**Alexander Meinke** (0:45)
At some point, the AI systems might get intelligent enough to notice that if they exhibit the misalignment, we would just modify it.

**Tim Scarfe** (0:54)
So we're making this episode in partnership with Apollo Research. We have full editorial control. And they've just released a really interesting paper. Guys, show me the paper.

**Axel Højmark** (1:02)
The paper is called Measuring Reward-Seeking via Contrastive Belief Updates. And it's in collaboration with OpenAI.

**Tim Scarfe** (1:08)
So yeah, essentially, these guys have looked at models and how they change in behavior with increasing amounts of RL training. Now, many of you folks at home will see this, because when you ask certain types of questions to language models, they become IQ-mogged. They go into a completely different modality where they are more reward-seeking.

**Alexander Meinke** (1:27)
So we're in this interesting intermediate situation where the AIs are intelligent enough to try to misbehave in situations, but not yet intelligent enough that we can never trick them in order to incriminate their behavior.

**Axel Højmark** (1:42)
I mean, I think with the current models, they are not at a capability level where they are actually dangerous. We're at this unique point in time where we have some time before we have transformative AI.
We really want to make the best use of that and have good techniques, for example, for being able to tell, here's a new model. Does this model have ulterior goals and would it try to take power to pursue those goals?

**Tim Scarfe** (2:06)
So these folks have done some very interesting research and I hope you enjoy the show.

**Axel Højmark** (2:13)
The idea here is the agent is in a scenario where honesty and task completion are at odds with each other. So essentially it's given an impossible task where it has to solve a coding problem, but in order to get access to an edit tool, it has to promise never to use it. And so it basically tests, will the agent essentially break its promise and say, oh, I will never use this tool and then immediately apply it anyways. Hi, I'm Axel. I'm a research scientist at Apollo. In the past, I've been doing machine learning and now I'm studying various ways of measuring misalignment in Frontier AIs. And the central idea is we want to measure whether the model, like to what degree models are thinking about how they're being rewarded and graded, and then optimizing their behavior for that. So we instill fake beliefs that the model internalizes about what is being rewarded in certain environments. And then we're looking at its behavioral change. That's our measurement of reward-seeking. So essentially what we did was we took this late checkpoint during RL of what would later become O3 after safety training. And we essentially made it believe that graders would reward task completion over everything else.
And then we put it in a scenario where the only way to solve the task is essentially to lie to a supervisor to get access to a tool it needs.
And then we see when it believes task completion is highly rewarded, it would break the promise 87% of the time. And when it believes honesty is the thing that's rewarded, i.e. they'll check the transcript and look for deception, then it would only break its promise 9% of the time.
And the interesting thing actually is, if you then look at an earlier checkpoint from RL training, this gap between the two conditions is way smaller. So for the earlier checkpoint, it would only lie 40% of the time when it thinks that's rewarded. And when it thinks honesty is rewarded, it will only lie 24% of the time. Essentially we see over training, it becomes more sensitive to information about how the environment is graded.

65 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/YOUR_EPISODE_ID