**Neev Parikh** (0:00)
METR, as you said, model of evaluation and threat research. The overall goal is effectively to try and measure catastrophic risk in a very scientifically rigorous way, have the ability to really get a handle on the kinds of risks that AI models are very likely to pose to us, be able to measure that really accurately, precisely. You want your tusks to have this, what we call, high ceiling effectively. Even at the very top end, there's still a lot of room as much as possible to keep improving your score, right? It's somewhat less useful if your task can just be maxed out at some point and then there's just no more improvement. There was no special prompting or anything where we're like, huh, that's cheeky. It wasn't that clever. It was somewhat clever, but not super clever where it was some subtle backdoor or anything like that. It was just like, oh yeah, I will just not train the model. Change the reference model and just copy it over so it will meet all the criteria of the task, but then it will be like zero training time. There are caveat thought there, but it was definitely interesting to see this kind of in the wild and completely unexpected. We weren't doing anything related to deception or something.
**Nathan Labenz** (1:04)
Hello, and welcome back to The Cognitive Revolution. Today, my guest is Neev Parikh, member of the technical staff at METR, or the Model Evaluation and Threat Research Organization. METR recently released a fascinating new benchmark for evaluating AI systems called Research Engineering Bench, or REBench for short, designed to assess how well AI agents can perform real machine learning research engineering tasks. The benchmark consists of seven challenging tasks across three categories, optimizing runtimes for performance, minimizing loss functions, and improving model win rates. To succeed, models have to do things like optimize GPU kernels, diagnose and fix corrupt models, and fine-tune language models for question-answering.
What makes this eval framework particularly interesting to me is how it approaches the challenges of comparing human and AI performance. Rather than using multiple-choice questions or other simply structured problems that might quickly saturate, REBench tasks are open-ended. They require experimental trial and error, and they're scored in such a way that allows for incremental progress with extra effort. The results show that leading models like Claude 3.5 Sonnet and OpenAI's O1 perform somewhere between the 10th and 40th percentile as compared to professional human machine learning researcher baselines, at least over an 8-hour time horizon. Interestingly, extending the AI's time budget by running multiple independent trials and then taking the best result significantly improved AI's relative performance, though still not to the level of top human experts.
Beyond the specific findings, I think this work is worth studying for several big-picture reasons. First, it represents a new class of AI evaluation, designed to push models out of their comfort zones. It requires reasoning over unfamiliar and in some cases quite unusual problems, effective use of tools, and the ability to maintain coherent plans over an extended period. These tasks simply cannot be solved through simple pattern matching or the regurgitation of training data. Second, the conceptual challenges the METR team faced in creating fair comparisons between humans and AIs highlight just how alien these systems really are. Humans need time to orient themselves to any given task, and accomplish little in the first two hours, but are much more able to continue making progress hour after hour. Whereas, by comparison, the AIs make progress almost immediately, but later on tend to get stuck in loops. Third, and perhaps most importantly, while current models still lag human experts, we are rapidly approaching capability thresholds that would enable significant automation of AI R&D itself, a scenario which for many years has been thought to signal the beginning of an intelligence explosion. Importantly, the METR team emphasizes that these results reflect relatively limited effort to optimize the AI agent's performance, which means we should expect better results with improved prompting and scaffolding. And of course, we now know that core model progress isn't stopping either. We recorded this episode on December 12th. Just eight days later, OpenAI introduced their new O3 model, which has once again made stunning progress on some of the hardest benchmarks ever devised. I fully expect it will move the needle on REBench 2, though by exactly how much, we'll have to wait and see. My conversation with Neev covers all this and more, from the nitty-gritty details of how the benchmarks work, to the surprising and in the context of such rapid progress on AI reasoning, I think quite chilling observations of reward hacking behavior. To finally, how we should understand the trajectory of AI development overall. As always, if you're finding value in the show, we'd appreciate it if you'd share it online, we'd love a review on Apple or Spotify, and we always enjoy your comments on YouTube. We invite your feedback too, either via our website cognitiverevolution.ai, where you can still submit questions for our upcoming AMA episode, or by DMing me on your favorite social network anytime. Now for a detailed look at the cutting edge of AI evaluation, I hope you enjoy this conversation with Neev Parikh from METR. Neev Parikh, member of the technical staff at METR, welcome to The Cognitive Revolution.
99 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000681243230