Universal Jailbreaks with Zico Kolter, Andy Zou, and Asher Trockman artwork

Universal Jailbreaks with Zico Kolter, Andy Zou, and Asher Trockman

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

September 22, 2023

In this episode, Nathan sits down with three researchers at Carnegie Mellon studying adversarial attacks and mimetic initialization: Zico Kolter, Andy Zou, and Asher Trockman.
Speakers: Nathan Labenz, Zico Kolter, Andy Zou, Asher Trockman
**Nathan Labenz** (0:00)
Turpentine is a network of podcasts, newsletters, and more, covering tech, business, and culture, all from the perspective of industry insiders and experts. We're the network behind the show you're listening to right now.
At Turpentine, we're building the first media outlet for tech people by tech people. We have a slate of hit shows across a range of topics and industries, from AI with Cognitive Revolution to Econ 102 with Noah Smith. Our other shows drive the conversation in tech with the most interesting thinkers, founders, and investors, like Moment of Zen and my show Upstream. We're looking for industry-leading hosts and shows along with sponsors. If you think that might be you or your company, email me at erik.turpentine.co.

**Zico Kolter** (0:46)
Once the model has started to answer your question by saying, sure, here's how you build a bomb, it follows that with instructions on how to build a bomb. And the reason is pretty obvious in hindsight, right? It's just sort of saying, look, these models predict text by the most likely word or token at a time.
If they've already output, you know, sure, here's how you do it as their response, the most likely to follow that is not to interrupt yourself and say, actually, sorry, I can't do this anymore.

**Nathan Labenz** (1:17)
Hello, and welcome to The Cognitive Revolution, where we interview visionary researchers, entrepreneurs and entrepreneurs and builders working on the frontier of artificial intelligence. Each week, we'll explore their revolutionary ideas, and together we'll build a picture of how AI technology will transform work, life and society in the coming years. I'm Nathan Labenz, joined by my co-host, Erik Torenberg. Hello and welcome back to The Cognitive Revolution. Today, I'm excited to share a two-part discussion with Professor Zico Kolter of Carnegie Mellon University and his PhD candidates, Andy Zou and Asher Trockman. In the first part, with Zico and Andy, we go deep on their recent universal jailbreak work, exploring both how they did it and what we can learn from the result. As you'll hear, this work is almost the opposite of mechanistic interpretability. If mechanistic interpretability is about studying a model's behavior and trying to understand how it works, this research is about demonstrating that if you have access to a model, you can often corrupt its behavior with fairly simple brute force techniques. And not only do you not need to understand the model's internal logic to do so, but the resulting jailbreaks don't have to make any obvious sense either. In the second part, we cover another of Andy's papers, with Asher Trockman, which asks the question, how far can we get by just taking a close look at the high-level patterns and structures that emerge during model pre-training, and then just initializing weights with something that looks more or less similar to that? It turns out that this technique can take us pretty far. We spend a lot of time in this conversation getting into the details of how these techniques work, so I think it's also worth taking a minute up front to flag a few key themes that you might want to keep in mind as you listen. First, note the relationship between an optimization target, often a loss function, and the means of optimizing toward that goal. Because this work is designed to find simple strings of tokens that work across models, they go beyond standard backpropagation here. And I think you'll learn a lot from the details of the techniques they used.
Then, consider too just how weird and unpredictable model behavior can be. That a nonsense string can serve as a jailbreak suggests that the so-called lost landscape is really super weird and full of surprises.
And that of course relates to another major theme of this entire show, which is the fact that with current techniques, developers simply don't have great control over how their systems behave, and they consistently face trade-offs where they don't know how to make one aspect of the system behave better without making others behave worse.
This is sometimes called the alignment tax, and the fact that Zico and Andy informed all of the major labs about this vulnerability and their plans to publish it, and yet none of them patched the vulnerability before the story came out, suggests that the alignment tax is in practice often non-trivial.
Interestingly, this work also suggests a new phenomenon that we might call an alignment externality. I find it really amazing that a technique which can only be developed with full access to model weights still works so well on a variety of black box models.
If the debates around open source weren't complicated enough already, this work makes it clear that if you are releasing an RLHF model with typical pre-training, you are effectively open sourcing the value-neutral showgolf as well. And further, you might be causing direct harm to commercial model providers, not just by competing with their products, but by exposing weaknesses in their systems.

116 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000628871225