Leading Indicators of AI Danger: Owain Evans on Situational Awareness & Out-of-Context Reasoning, from The Inside View artwork

Leading Indicators of AI Danger: Owain Evans on Situational Awareness & Out-of-Context Reasoning, from The Inside View

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

October 16, 2024

In this special crossover episode of The Cognitive Revolution, Nathan introduces a conversation from The Inside View featuring Owain Evans, AI alignment researcher at UC Berkeley's Center for Human Compatible AI.
Speakers: Nathan Labenz, Owain Evans, Michael Trazzi
**Nathan Labenz** (0:00)
Hello, and welcome to The Cognitive Revolution, where we interview visionary researchers, entrepreneurs, and builders working on the frontier of artificial intelligence. Each week, we'll explore their revolutionary ideas, and together, we'll build a picture of how AI technology will transform work, life, and society in the coming years. I'm Nathan Labenz, joined by my co-host, Erik Torenberg. As a developer, the journey from concept to production-ready large language model apps is fraught with challenges. Dealing with unpredictable language model outputs, hallucinations, and ballooning API costs can all be blockers to shipping your next AI-powered feature. That's where Advanced RAG comes in. With the new RAG++ course from Weights and Biases, you can overcome these hurdles and build reliable production-ready RAG applications. Go beyond proof-of-concept and learn how to evaluate systematically. Use hybrid search correctly and give your RAG system access to tool calling. Based on 21 months of running a customer support bot in production, industry experts at Weights and Biases, Cohere, and Weaviate show you how to get to a deployment-grade RAG application. This offer includes free credits from Cohere to get you started. Make real progress on your large language model development and visit wnb.me.cr to get started with their RAG++ course today. That's wnb.me.cr to get started with their RAG++ course today.
Hello, and welcome back to The Cognitive Revolution. Today, I'm excited to share a special crossover episode from The Inside View, featuring a conversation on situational awareness, out-of-context reasoning, and other AI safety topics between Owain Evans, AI alignment researcher at the Center for Human Compatible AI at UC Berkeley, and creator and host of The Inside View, Michael Trazzi. Owain is someone I've known casually through friends for many years, and it's been amazing to see him develop into such a prolific and influential researcher. His Google Scholar page lists 12 papers just since 2022, most of which serve to carefully map out some dark, but perhaps important corner of large language model capability space, and some of which you've likely heard of, including The Reversal Curse, which showed that LLMs trained on information like A is B often fail to learn that B is A, and also Connecting the Dots, which is discussed in this episode, and which shows that at least to some degree, large language models are capable of inferring censored information from the implicit hints contained in their training data. In this episode, Owain explains why situational awareness matters, particularly in the context of deceptive AI scenarios, and discusses his research on measuring situational awareness, including the development of a benchmark to assess this capability. This topic has honestly never felt more relevant to me, because I've been coding with the new O1 model quite a bit over the last few weeks, and while to be honest, I have to admit that the model is in most respects more cognitively capable than I am, its situational awareness still seems rather weak, making this both one of a shrinking set of dimensions in which humans still have a meaningful advantage over the AIs, and one worth watching very closely as AI systems continue to evolve. I've been wanting to have Owain on the show since I started it, and I look forward to covering more of his work in a future episode. For now, if you're finding value in the show, we of course appreciate it when folks take a moment to share it with friends, or to write an online review on Apple Podcasts or Spotify. We always welcome your feedback via our website, cognitiverevolution.ai, or by DMing me on your favorite social network. Finally, before getting started, I really encourage you to check out The Inside View, either on YouTube or at theinsideview.ai for more of Michael's content. I specifically recommend his episode on Anthropics Research into Sleeper Agents with Evan Hubinger. I'm also looking forward to his upcoming documentary on SB 1047, which will feature a number of past Cognitive Revolution guests, including Dean Ball, Timothy B. Lee, Leonard Tang, Dan Hendricks, Nathan Calvin, and Flo Cravello. Now, let's get into the weeds on the critical topic of situational awareness in AI systems, with alignment researcher, Owain Evans, and Michael Trazzi, creator of The Inside View.

**Owain Evans** (4:12)
If models can do as well or better than humans, who are like AI experts, who know the whole setup, who are like trying to do well on this task, and also they're doing well like on all the tasks, including like some of these very hard ones, I think that would be like one piece of evidence, right? Where you could say, look, two years ago or something, right? Models were not, were like way below human level on this task.
Now, they're like above human level. There's evidence here that they have the kind of skills necessary to understand when they're being evaluated, to take actions that go against their training data. So this would be, I think, yeah, a piece of evidence where you could say, look, given this performance, we should think carefully about alignment of the model, like what evidence we have for alignment. We should maybe try and understand the skills. It's like, how is the model doing so well? Is it memorization or specialized fine tuning? This would be a reason to try and find out, okay, how general is this skill? because I think it would certainly be concerning if models were getting plus 90% on this.

111 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000673311647