**Erik Torenberg** (0:02)
Before we dive into today's episode, I want to tell you about a new show from Turpentine called Modern Relationships. On the season ahead, I sit down with power couples in tech and leading relationship thinkers to explore how ambitious people actually make partnerships work. Whether you're dating in a relationship or just curious how technology is reshaping modern love, I think you'd enjoy this on your feed. Our first episode features founders funds Delian Asparouhov and tech researcher Nadia Asparouhov, who take us through their evolution from dating to marriage to parenthood. With absolutely no filter on the challenges and growth along the way. You can find Modern Relationships wherever you get your podcasts. Now, on to today's episode.
**Scott Emmons** (0:37)
If the model is doing some highly capable, sophisticated behavior, like if it is implanting a quite sophisticated backdoor in your code, this doesn't happen by accident. The hope is that the language model activations already have this lower dimensional or this higher level representation of the sophisticated behavior that we're seeing.
**Eric Jenner** (0:57)
Someone might poison your training data. So like all these models are trained by just collecting data from the open internet. So someone can just upload something to a website and then eventually it's going to be scraped and end up in the training data.
**Luke Bailey** (1:08)
The behavior that the attacker is trying to elicit from the model may actually be high loss to begin with.
**Eric Jenner** (1:14)
The more of reasoning moves away from natural language into model latents, I think the harder it gets to monitor models. And I do think that's one of the scariest things going on right now from a safety perspective.
**Nathan Labenz** (1:23)
Hello, and welcome back to The Cognitive Revolution. Today I'm speaking with Luke Bailey, Eric Jenner, and Scott Emmons, who along with their fellow lead authors Alex Serrano, Abhay Shashadri, Mikhail Soleznyov, and Jordan Taylor have recently published a fascinating new paper called Obfuscated Activations Bypass Large-Language Model Latent-Based Defenses. Latent-based defenses, regular listeners will know, include a number of techniques for classifying and in some cases interpreting large-language models' internal activation patterns with the hope of detecting and preventing harmful or otherwise undesired behaviors. To many observers, myself included, these defenses have seemed like a particularly promising possible solution to AI safety and control problems, with at least some hope of achieving adversarial robustness, since presumably it would be hard to perform specific behaviors without passing through their natural cognitive precursor states. However, through an extensive series of experiments targeting three major classes of latent-space defenses, the authors demonstrate that it very often is possible, with a mix of methods that they collectively call Obfuscation Attacks, to elicit specific target behaviors while shifting internal activations so as to avoid detection. For example, for supervised probes, which learn from human-labeled examples to identify the activation patterns associated with harmful inputs or outputs, they show how to maintain the specific target behavior while shifting activations so as to avoid detection. Similarly, for out-of-distribution detectors, which use statistical techniques to identify deviations from normal activity and have been used to detect backdoor attacks like Anthropic's famous sleeper agents, they again develop methods that preserve the unwanted behavior while keeping activations within expected statistical bounds. And for sparse autoencoders, which of course can be used to understand the high-level concepts that a model is working with at any given layer and token position, they again demonstrate that it is possible to produce a specific desired or undesired behavior while avoiding detection at any specific layer in the model. Though, performance does gradually degrade when they try to evade SAE detection at multiple layers. The amount of work and the number of experiments in this paper is remarkable, and there is a ton of detail in this conversation. We cover multiple different kinds of attacks, including prompt suffixes, soft prompts, and data poisoning, the structure of the lost functions that they optimized against, the level of access required for these attacks to work, and the fact that some of these attacks work even when the defender moves last, some really fascinating cat-and-mouse type experiments that show attacks working over and over again over 70 rounds of back and forth between attack and defense, and even their attempts to make models easier to monitor through an approach that they call Obfuscated Adversarial Training, which they found does make it harder, but still by no means impossible, to perform these obfuscated activation attacks.
Overall, while these findings clearly raise important questions about how effective interpretability-adjacent approaches to AI safety and control will be, and again suggest that defense in depth may really be all we have, I'm still not quite sure how best to understand what's really going on here. As you'll hear, the authors believe that these results show that models can represent the same concepts in many different ways, and that might prove to be right. But there is enough evidence for consistency and fragility in representations, and here I'm thinking about recent papers like the Platonic Representation Hypothesis, Mielnanda and team's refusal in large language models is mediated by a single direction, and most recently the superweight in large language models, that for now I'm keeping an open mind and planning several more episodes in this area. Including another interpretability update with GoodFire co-founders Dan Balsam and Tom McGrath, and hopefully soon the conversation with legendary adversarial defense buster, Nicholas Carlini from DeepMind. For now, if you find this sort of exploration of AI internals valuable, we'd appreciate it if you'd share it with a friend, write a review on Apple Podcasts or Spotify, or leave us a comment on YouTube. And of course, we always welcome your feedback and suggestions. You can drop us a note via our website, cognitiverevolution.ai, or by DMing me on your favorite social network. With that, I hope you enjoy this technical deep dive into Latent-Space Defenses and the attacks that for now defeat them. With Luke Bailey, Eric Jenner, and Scott Emmons. Luke Bailey, Eric Jenner, Scott Emmons, authors of the new paper Obfuscated Activations Bypass Large Language Model Latent-Space Defenses. Welcome to The Cognitive Revolution.
119 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000684496863