**Nathan Labenz** (0:00)
Hello, and welcome back to The Cognitive Revolution. The Cognitive Revolution is brought to you in part by Granola. If you are a regular listener, you've heard me describe the blind spot finder recipe that I'm using to look back at recent calls and help me identify angles and issues I might be neglecting. But it's also worth talking about how Granola can help raise your team's level of execution by supporting follow-through on a day-to-day basis. This past week, for example, I had several working sessions with teammates, and I committed to a number of things. In the past, to be honest, there's a good chance I'd have forgotten at least a couple of the things I said I'd do. But with Granola, I can easily run a to-do finder recipe and get a comprehensive list of everything I owe my teammates. This is the sort of bread and butter use case that has driven Granola's growth and inspired investment from execution-obsessed CEOs, including past guests, Guilherme Rauch of Vercel and Amjad Massad of Reblet. See the link in our show notes to try my blind spot finder recipe and explore all of the ways that Granola can make your raw meeting notes awesome. Now, today, my guest is Geoffrey Irving, a pioneering machine learning researcher who's co-authored seminal papers with a who's who of giants in the field and who is now chief scientist at the UK AI Security Institute, which is in all likelihood the most situationally aware government entity in the world today. With roughly 100 technical experts on staff and a mandate that includes threat modeling, pre-release frontier model evaluation for dangerous capabilities spanning biosecurity, cybersecurity, and loss of control, advising the UK government on strategies to reduce catastrophic risk, funding independent frontier research, and engaging in global diplomacy, Geoffrey has one of the most broad and commanding views of the AI landscape that you'll find anywhere. And while he is optimistic about our ability, in the fullness of time, to solve the major open problems in AI safety, for today, without a hint of hype, he paints a genuinely alarming picture. Our theoretical understanding of machine learning is nascent. Nobody, he argues, should be particularly confident in their mental models of how AI will go. Models already outperform a majority of experts on a great many security-related tasks, and there is no good reason to expect that their progress will stall. Reinforcement learning is working well beyond strictly verifiable tasks, and jaggedness matters much less when even the model's weak spots are as good or better than the best humans. The many increasingly sophisticated bad behaviors we've seen over the last 18 months are broadly all different versions of reward hacking, a problem for which we lack theoretical or practical solutions. As such, we likely won't get that many nines of reliability from current safety techniques, and there is some reason to expect that they could all fail at the same time for the same reasons. It is getting harder to jailbreak models, but the AC Red team has never failed to do so. And meanwhile, eval awareness is an open and growing problem. Voluntary cooperation between frontier model developers and the AC is working pretty well, but not everyone is participating. The AC, for its part, is seeking to fund theoretical research in areas like information theory, complexity theory, and game theory, which might produce stronger guarantees. But these fields, like most of the rest of the world, are just beginning to take AI seriously at all. Geoffrey is an intellectual powerhouse, but I came away from this conversation just as impressed with the UK AC as a whole. This is an organization staffed with top-notch talent that has its finger on the pulse of industry development, and is speaking very accurately and plainly about AI's trajectory and how many major questions remain unanswered, even as frontier model company CEOs tell us that they are less than three years away from creating expert-level AI machine learning researchers. With that, I hope you are focused and motivated by this conversation about the AI state of play with Geoffrey Irving, Chief Scientist at the UK AI Security Institute.
Welcome to The Cognitive Revolution.
**Geoffrey Irving** (4:16)
Thank you. I'm excited to be here.
**Nathan Labenz** (4:18)
I'm excited for the conversation. We've exchanged messages for a while and have been building up for this, and I'm excited that the moment is finally here. And you have really a storied publication history that goes back to working on the original TensorFlow papers with some guy named Jeff Dean, being a co-author on the original RLHF paper, working on concepts years ago.
**Geoffrey Irving** (4:43)
For language.
**Nathan Labenz** (4:46)
A caveat, but still, right there alongside Paul Cristiano, some early AI safety papers with no less than Dario on concepts of using debate to try to bootstrap into a stable equilibria and stable AI safety regimes, and even published a call for social scientists to enter the field of AI safety with one Amanda Askell. I would be very interested to hear how it was that you came to have such a good nose for where AI was going so early on. All these things are well before Chach-Upt.
122 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000752309675