**Joel Becker** (0:00)
So, Meta stands for M-E-T-R. The first two letters, model, evaluation, that is, we think about what the capabilities of AI models might look like today and tomorrow, as well as their propensities, what they'll actually do in the wild, given that they have some level of capability. And then threat research is the final two letters. We try to connect those capabilities and propensities to particular threat models that we have in order to determine whether AI models pose enormous or catastrophic risks to society. So, the secret, if you read this article about how I became the number one most profitable trader on Manifold, mostly comes down to this one market where...
**Alessio** (0:39)
Hey, everyone. Welcome to the Latent Space Podcast. This is Alessio from the Recurrental Labs, and I'm joined by Swips, editor of Latent Space.
**Swips** (0:45)
Hello, hello. We're back in the studio with Joel Becker from METR. Welcome.
**Joel Becker** (0:49)
Thank you very much, guys. It's a great pleasure to be here.
**Swips** (0:51)
So, Joel, your work has impacted the AI field a lot, especially over the last year. I invited you for the AIE Summit, which thank you for speaking as well and doing the workshop. And you have a lot of papers that have been very impactful. But I guess upfront, a lot of people, like METR just burst onto the scene. Could you explain and introduce METR?
**Joel Becker** (1:11)
Yes. So METR stands for M-E-T-R. First two letters, model, evaluation. That is, we think about what their capabilities of AI models might look like today and tomorrow, as well as their propensities, what they'll actually do in the wild, given that they have some level of capability. And then threat research is the final two letters. We try to connect those capabilities and propensities to particular threat models that we have in order to determine whether AI models pose enormous or catastrophic risks to society.
**Swips** (1:39)
Yeah. Would you say that you've done a lot more M-E and T-R is like the next phase, or is there a T-R side of work that I'm less conscious of?
**Joel Becker** (1:46)
I think there's some T-Rs. Some of the most publicized work does look more like VME. It looks like this time horizon stuff and the developer productivity RCTs, stuff like that. But there's this wonderful report on our website, GPT-5 report and analogous one for GPT-51 as well. Trying to make this more structured case that it doesn't pose these really large scale risks, eventually coming to the conclusion that it doesn't. But it's worth thinking, why exactly is that the case? If you and I work with GPT-5, it does seem very capable that matches up to benchmark scores. Why is it not able to do something really enormously wrong? We go through the evidence, we find we think it's not capable enough, on the basis of some of this capabilities evidence that you've alluded to, to commit these catastrophic harms, and it's not going to be able to do this. But perhaps in future, we'll think it's capable of doing pretty extraordinary things, kinds of things that would be necessary to provide really serious threats. And then maybe you'd lean more on the propensities part. Are the protections that we have against these dangerous capabilities sufficient for it not to pose an existential threat? That sort of thing. So I think it's, I think threat research very much is there, very much is something that we're aspiring towards. In some ways, you might, sorry, see the capabilities evidence as a kind of input.
**Swips** (2:52)
Yeah.
**Alessio** (2:53)
Have the threat models been updated a lot? Or do you feel like you're still using the same threat models as GPT-2 or Paperclip Factory, blah, blah, blah? Or like, how much are you increasing the bar?
**Joel Becker** (3:05)
Yeah, so I'm not an expert in the threat modeling piece, more in the capabilities piece. I do think they've been changing to some extent. So something like the autonomous replication threat model, that is being able to set yourself up and control resources, something like that, has been de-prioritized relative to R&D acceleration. That is the possibility there could be some capabilities explosion inside of a lab, and that could be destabilizing for all sorts of reasons that we could talk about. So mainly we're focusing on that latter one, although we do think about a number of threat models.
**Alessio** (3:33)
Yeah, let's talk about the ME side. So I would say the model time horizon chart is probably the most quoted, I would say, both in investment decks that I see and just general on Twitter. What was the origin story of it? And any other color you want to give on it to introduce it to the audience?
57 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000751972505