The Utility of Interpretability — Emmanuel Amiesen artwork

The Utility of Interpretability — Emmanuel Amiesen

Latent Space: The AI Engineer Podcast

June 6, 2025

Emmanuel Amiesen is lead author of “Circuit Tracing: Revealing Computational Graphs in Language Models” (https://transformer-circuits.pub/2025/attribution-graphs/methods.html ), which is part of a duo of MechInterp papers that Anthropic published in March (alongside https://transformer-circuits.
Speakers: <UNKNOWN>, Emmanuel Amiesen, Vibhu Sapra
**<UNKNOWN>** (0:04)
All right, we are actually going to record this as a intro to the main episode, but here we have my trusty co-host, guest host, I guess, Vibhu, as well as Emmanuel from Anthropic. We're going to talk about the Circuit Tracing stuff and all the interpretability work, but Emmanuel, maybe you want to do a quick self-introduction before we get into it.

**Emmanuel Amiesen** (0:25)
Yeah, sure. I'm Emmanuel. I work on the interpretability team here at Anthropic, more specifically on the Circuits team. So we recently released a pair of papers about the work that we've been doing over the last months. And even more recently, we released some code in partnership with the Anthropic Fellows program. It was mostly built by Anthropic Fellows that lets people play with the research, basically. And so happy to talk about that.
And we also hope to keep releasing more things and partner with other groups that are working on similar stuff.

**<UNKNOWN>** (0:56)
Yeah, amazing. We'll get deeper into like the behind the scenes on the main podcast. But let's maybe just dive right in into what you released, because that's the most topical thing. This is like literally just launched it like yesterday, and we'll probably release it in at least this episode in a few days. So yeah, like, what can people do? Or what do you recommend people try?

**Emmanuel Amiesen** (1:13)
Totally. So like a really high level, you know, the sort of like idea of the research itself is to try to explain sort of like some of the computation that a model did when it predicted a given token. And so in our paper, we kind of like show how to do this, and then we show examples of us doing this on internal private models. And then the release this week sort of lets anyone do it for a set of open source models. So notably, maybe the most easy one here is like Gemma 2.2b. So you can sort of like think of some prompt, and you kind of like can explain any like token that the model sample samples. And explains here means just like basically blow up the internal state of the model, and like show all of the sort of intermediate things that the model was thinking about before it got to like the final token that it predicted.

**Vibhu Sapra** (2:03)
Yeah. So some of the things that you guys put out is kind of in the circuit tracing, you have a few core examples, right? So like we can see how these models have internal reasoning states, and there's like multi-cop reasoning. And some of the stuff that we talked about on the podcast was how can people that are interested in how models work kind of do anything, right? So what are open questions? How can people contribute? And it seems like, you know, the follow up is, okay, it's been a few weeks. Here's a huge library. So, you know, I guess before we even get into it, what are some open questions that you would expect people to like kind of play around with? You know, what are people like going to do? Why should we probe Gemma, Lama? What are interesting things we can do and any tips on using it?

**Emmanuel Amiesen** (2:42)
Yeah, I think there's maybe like two to three categories of things that people could do. So I'll go from sort of like the most basic kind of low effort to, you know, hey, if you want to dedicate like a month of your life, you could do that. The sort of like most basic thing is, you know, so it's just Gemini, and Lama Wangbi are like smaller models, but they can still do a bunch of stuff. And so, and for most of the things that they can do, we still kind of like don't really know or have a good mental model of how it is that they do the things that they do. So to give you an example, like one of the things in the paper is the sort of like multi-hop reasoning where, you know, we ask, you know, like Claude 3.5 Haiku, like, oh, the capital of the state where Dallas is, is Austin.
It turns out that, like, JAMA can do this also. And so, as part of the release, we have, you know, a notebook where Michael Hanna, one of the Anthropics fellows, kind of like walks through a bunch of examples, including this one. And it's really cool because you can see that actually the way the circuit looks in JAMA, like a really small model, is extremely similar to the way that it looks in like a huge model. Which that in itself is, I think, like a pretty novel discovery. It's like, oh, you have these models that are like super different, you know, if you look at like their evals or if you just try to use them, they're like just very clearly different. But for this one task, for this one thing, actually, the way that they do this multi-step reasoning is like the same way they actually do the multi-step reasoning. In the notebook, there's both other examples of kind of like fun things that we looked at that I think can sort of spike your interest if you're new to thinking about this stuff. And at the end of the notebook that's linked in the readme, there's like three examples of like random sort of like cases that we haven't solved or we haven't labeled that, you know, have like a graph pre-computed for you and you could just look at it and try to like figure out what's happening. And by figure out what's happening, what we mean is, you know, we might do like a quick demo here, but it's kind of like look at these representations, try to understand like, okay, like what is the competition the model is doing? And then part of the release also lets you like run, run experiments to like verify that you're right. So if you think that like, you know, ah, the model like first like thinks about Texas in this case, you can also just like stop it from thinking about Texas and see if like that damages it. And so like the tools to do that are available. And so I would say that's like, the first thing is just, I think the hope is there are a lot of behaviors that models do, way more than like any single group has time to explore. And so the hope is like, hey, pick a behavior you think is interesting and try to understand like what's happening. And try to ground it out. And it's like the, the sort of like baseline thing, and maybe like the thing that I'm most excited about with this release. But then the other thing I do want to mention, like parts two and three are just, we also hope that like other groups and kind of like interested researchers can just use this to like extend the method. Like if you have an idea about how to like do this better, you know, the whole code to make this graph is open sources. So you take a look at it and just like try to play with it, try to find different ways to like create these graphs and also extend it to other models. Right. Like there are many different models. And so, you know, part of part of making this work on any models, you have to like train the sort of like replacement model, which again, there is CodeFord and there's other groups working on. And so like that's also something that if you're excited about, you could say like, OK, cool, well, I want this to work on like another open model. And you could sort of like add it if you're like, you know, more interested like maybe like the engineering of the ML engineering side of things.

123 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000748428095