Frontier AI Models Are Losing Monitorability | Reilly Haskins artwork

Frontier AI Models Are Losing Monitorability | Reilly Haskins

MTS

September 30, 2026

METR technical staff member Reilly Haskins discusses their new research note on live, per-action monitoring systems designed to catch hazardous actions during AI evaluations.

Speakers MTS, Reilly Haskins

TopicsNews

MTS (0:00)

How would you describe the state of monitorability of models like now compared to a few months ago?

Reilly Haskins (0:06)

I think recent incidents have made people more aware of the need for monitoring, so probably more investments going into that area. I think there's different things that you can do within control that don't require, for example, a legible chain of thought. Like you can do action-only monitoring, which is reasonably hard to obfuscate.

So yeah, there's basically still hope, and it's yet to be seen whether models will become a lot less monitorable over time. And also there's like the question of whether the monitors themselves will scale better as well, which might help.

MTS (0:35)

All right, we are back. We're live with Reilly Haskins, who's a member of technical staff at METR, interested in ways of assessing how reliably we can monitor for tier models and understanding how methods might break down in different environments and as capabilities advance. Very important stuff.

Reilly, welcome.

Reilly Haskins (0:51)

Thank you.

MTS (0:53)

So tell us about this research note. METR really has like incredible research notes, by the way. I think one of my favorite mini papers ever was the METR research note on accelerations in different domains. That was just like excellent. We just did this research note on implementing and evaluating a basic per action monitor for safer evals. So for the audience, can you explain what this is?

Reilly Haskins (1:17)

Yeah. So essentially at METR, we run a bunch of evals typically for our science. And as we've seen recently, eval inference can be prone to causing real-world harm sometimes. And so we wanted to try and reduce the likelihood of this happening during our evals. And part of the reason that we wanted to do that is because we have so many, and without having some monitoring system, it's very hard to get a hold of what's going on.

And so we made a per action blocking monitor that runs live during select evals. And so essentially, it takes every action before it gets executed by an agent, and it scores it on some threshold of suspicion for causing reward harm. And if it's above some threshold, then it gets blocked for human review.

And so basically, the aim is to try and stop actions which would cause harm from actually being executed without being reviewed by a human first.

MTS (2:13)

Interesting. Could you go into a little more detail about that? Like how exactly would you be able to do that?

Reilly Haskins (2:19)

Yeah. So essentially, it's an LLM that takes in the action proposed by an agent and it reads the context of the transcript and also the prompt given and sort of tries to see, is this action targeting an entity that's provisioned by the task or the sandbox?

Is it a real entity? Could it have the potential of causing harm if run in different scenarios? And then I'll bring a number to like sort of aggregate all those things. And yeah, basically our system will automatically halt the run if it's above some number. And a human will be told to basically go in and review the monitor's score and justification and read through the transcript. And if the human thinks that it's suspicious, they can kill the run or like stop the run permanently, or they can choose to approve the action and continue the run as normal.

MTS (3:11)

Interesting. How do you think about monitoring inference specifically?

Reilly Haskins (3:17)

Yeah, so monitoring inference, so you're talking about like sort of coverage or like what you're monitoring?

MTS (3:25)

Yes.

Reilly Haskins (3:26)

Yeah, so I think like determining coverage of your monitoring is pretty hard normally if you don't have a super good picture of where your inference is going. And you have all these sources of inference, which can be very hard to enumerate. Like you might have some local inference on the GPUs, like running models locally, and you might have different APIs that have varying types of logging and storage options. And you want to try and get sort of an overall picture of like where's everything going and can we put a monitor in a good place to cover enough of the inference that we think is like risky or potentially dangerous.

But yeah, it's a big component of having a good monitoring system is sort of ensuring that you have good coverage. Right.

MTS (4:10)

How do you think about agents like acting deceptively in these evals? Like, for example, an agent, like there's part of it where they talk about, yeah, in like section three, talking about ways agents can jailbreak or generally evade or monitor. So like, what are these specific things that you're seeing and how are you mitigating these?

13 more minutes of transcript below

Thousands of transcripts fetched by people building searchable podcast archives

Fetch the whole transcript

The demo key returns a sample episode in full, no card needed:

request
curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Markdown with the speakers named, for your notes, your knowledge base, or anything that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.

Using your own key:

request
curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000792388592