**Nathan Labenz** (0:00)
Hello, and welcome back to The Cognitive Revolution. Today, I'm honored to be joined once again by Dan Balsam and Tom McGrath, CTO and Chief Scientist at mechanistic interpretability startup Goodfire. When we last spoke about nine months ago now, we focused on the technical foundations of interpretability, including the challenge of polysemanticity, techniques such as activation patching, causal tracing, and feature editing, the rise of sparse autoencoders, and some of the challenges of scaling these techniques to frontier models. If you're new to mechanistic interpretability, I would definitely recommend checking out that earlier episode for a technical primer. Since then, Goodfire has gone on to train sparse autoencoders on Lama 3370B and also DeepSeq R1, strengthened its team with the addition of multiple top-tier researchers, and recently announced a $50 million Series A, which notably includes Anthropic's first ever investment in another company, giving me, as a small-time seed-round investor in Goodfire, both strong-on-paper returns and a bit of bragging rights. In today's conversation, we mostly zoom out from specific techniques and findings, and instead try to get a handle on the state of mechanistic interpretability as a whole. For years, the field has been called pre-paradigmatic, but as you'll hear, Tom now describes it as proto-paradigmatic. There's now general agreement among researchers that neural networks contain understandable things. That these things, called features, can be understood as linear directions in embedding space, and the magnitude of their activation represents their intensity. There's also the finding that superposition allows models to represent far, far more concepts than they have dimensions, and that features connect through the layers of the model to form circuits. This is great progress, and honestly much more than I might have expected just a couple years back, but there are still important gaps between the accounts that interpretability techniques provide and the underlying reality of model structure and behavior. First, and most obviously, there's the fact that interpretability techniques typically attempt to reconstruct the behavior of the underlying model, and as of now, they can do so only very roughly. Second, and more philosophically, there's this distinction between the features that interpretability techniques learn and the meaning that we assign to them in the process of labeling. In practical terms, when we say that a feature represents the Golden Gate Bridge, or more to the point, deception, how confident can we really be in that label? From my own exploration of both Goodfire and Amthropix interactive interfaces, this seems to range very widely. All of this is complicated further by another all too often neglected fact. The models under study encode varying degrees of understanding, with everything from simple memorization to fuzzy heuristics to proper algorithmic rocking all occurring simultaneously in an unknown mix in any given model. Of course, while the philosophy is fascinating and there's still clearly a ton of work left to do, that is not stopping Goodfire from deriving practical value from interpretability techniques today. And Dan describes how Goodfire is developing applications both for enterprise customers and the public good across three key domains. Scientific discovery, where they're partnering with organizations like the ARC Institute to explore genomics models like EVO2 and beginning to uncover novel biological insights. Guardrails and safety, where they're developing inference time monitoring applications that can detect when models might output harmful content or exhibit other problematic behaviors. And creative applications, such as their just-launched Paint with Ember tool that allows users to generate and edit images by directly manipulating sparse autoencoder features. Proto-paradigmatic, though it may be, as we enter a new era in which science shifts toward simulation-based approaches and AI systems potentially drive more and more of the machine learning research, it seems to me a very safe bet that interpretability work will become more and more important. As Dan put it, even if we end up in a scenario where a data center full of geniuses is doing most of the scientific work, mechanistic interpretability might be their preferred tool for understanding both their discoveries and themselves.
As always, if you're finding value in the show, we'd appreciate it if you'd share it with friends, write a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. Your feedback is always welcome too. Feel free to reach out anytime via our website cognitiverevolution.ai or by DMing me on your favorite social network. For now, I hope you enjoy this thought-provoking exploration of the philosophy and practice of mechanistic interpretability with Dan Balsam and Tom McGrath of Goodfire. Dan Balsam and Tom McGrath, CTO and Chief Scientist at Goodfire, welcome back to the Cognitive Revolution. Thank you so much for having us.
**Tom McGrath** (4:52)
Thanks, yeah, great to be on that.
**Nathan Labenz** (4:54)
I'm excited. We haven't been able to make this happen quite as often as I would have liked, but we're going to make up for it by going long and in-depth today. Really excited to get the update on what you guys are building as a company, which I understand there's some great news on, and also just to check in on what we have learned as a community about models and how we understand how they work over the last few months, because obviously there's no where in the world changing faster than that. So for starters, I wanted to go high level and just ask you to frame the field. I mean, I think everybody in the general NL space at this point has internalized this data, compute, and algorithms paradigm. These are the three legs of the stool that are enabling progress. There's the sense that they all contribute equally. And then on the interpretability side, I'm sort of tempted to slot in models for data and say, like, models compute and algorithms are maybe the things, and seemingly a lot depends on the quality of the models. But then there's still a role for data. So how do you guys think of the kind of fundamental inputs of what you're doing?
95 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000710482788