Popular Mechanistic Interpretability: Goodfire Lights the Way to AI Safety artwork

Popular Mechanistic Interpretability: Goodfire Lights the Way to AI Safety

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

August 17, 2024

Nathan explores the cutting-edge field of mechanistic interpretability with Dan Balsam and Tom McGrath, co-founders of Goodfire.
Speakers: Nathan Labenz, Tom McGrath
**Nathan Labenz** (0:00)
Hello, and welcome to The Cognitive Revolution, where we interview visionary researchers, entrepreneurs, and builders working on the frontier of artificial intelligence. Each week, we'll explore their revolutionary ideas, and together, we'll build a picture of how AI technology will transform work, life, and society in the coming years. I'm Nathan Labenz, joined by my co-host, Erik Torenberg. Hello, and welcome back to The Cognitive Revolution. Today, I'm thrilled to be speaking with Dan Balsam and Tom McGrath, co-founders of Goodfire, a new company focused on mechanistic interpretability of AI models. Dan serves as CTO, bringing his experience as a serial startup engineer, while Tom, the chief scientist, comes from a background in AI safety research at DeepMind. As you know, mechanistic interpretability is the nascent science that seeks to understand AI models in our workings and explain why they do what they do, and which promises to light the way to engineering solutions to problems of AI control and safety. Like many subfields of AI, mechanistic interpretability has delivered incredible progress over the last couple of years as all three leading developers, led by Amthropic, but with DeepMind and OpenAI right on their heels, as well as Academia, which has been enabled by Lama3 and other open source foundation models, have made real progress on the great AI black box problem. Most recently, sparse autoencoders have emerged as a powerful tool for isolating and studying the concepts, also known as features, that large language models are learning and which are already beginning to be used for monitoring LLM internal states and steering their behavior. If you played with Golden Gate Clawed, where Amthropic intervened to set the Golden Gate Bridge feature to an artificially high level, or just had a laugh at some of the examples posted online, you already have an intuitive sense for what these sparse autoencoder discovered features can do.
In this conversation, which I hope will be just the first of many interpretability roundups with Dan and Tom, we cover a wide range of interpretability topics, including the concept of polysemanticity in neural networks and the challenges that this phenomenon poses for interpretability, engineering techniques for large-scale interpretability studies including activation patching, causal tracing, and feature editing, the rise of sparse autoencoders including their architecture, training processes, and outputs, early progress in autointerpretability which is the use of language models to label the features isolated by sparse autoencoders, the potential for interpretability to extract scientific knowledge from domain-specific models such as those trained on protein folding, weather prediction, and other problems, the technical challenges of serving interpretable models in production environments at scale, and how new AI architectures may impact interpretability science going forward. Toward the end, Dan and Tom shared their vision for Goodfire, which recently raised $7 million in seed funding, led by Lightspeed Ventures, to develop and ultimately productize this line of research for open-source models. The goal is to empower developers, and even everyday users, to look inside models, diagnose issues, and intervene in ways that improve performance and reliability, and ultimately to create a world in which companies simply don't deploy AI models that they don't understand. I am personally very excited about this direction, and I appreciate that the Goodfire team has allowed me to invest a very small amount in the company as part of this fundraising round. As always, if you're finding value in the show, we'd appreciate it if you'd share it with friends, post online, or leave us a review on Apple or Spotify. I love hearing from listeners and strive to respond to all messages, so please do feel free to DM me on your favorite social network anytime. Now, let's dive into the fascinating and fast-moving world of AI interpretability with Dan Balsam and Tom McGrath, co-founders of the new mechanistic interpretability startup, Goodfire. Dan Balsam and Tom McGrath, CTO and Chief Scientist of the new mechanistic interpretability company, Goodfire, welcome to The Cognitive Revolution. Thank you for having us.

**Tom McGrath** (4:03)
Great, thanks. Yeah, we're really excited to be here.

**Nathan Labenz** (4:06)
I'm excited for this too. We've been talking offline for the last couple of months as you guys have been putting this company together and I was excited about it enough to make a very small investment. So a disclaimer for this episode is that I am an interested party in this company. Although for anyone else who might seek my investment, just know that it would be an extremely tiny check size. But love what you guys are aspiring to do here in terms of figuring out how to not just get deeper into mechanistic interpretability, but to figure out how to make a product of that and to make that something that the whole world can get access to and take advantage of and collaboratively explore together. So I'm really excited to unpack this in detail and get your take on where we are in interpretability. I think this is going to be a lot of fun. For starters, you guys want to just introduce yourselves. I think you have quite different backgrounds, which is interesting. And people who are maybe inclined to be intimidated by mechanistic interpretability might see some inspiration, especially in your background, Dan, but give us both of them. Yeah. So I have been a serial early employee at startups, serial founding engineer for most of my career. Most recently at a Series B AI recruiting startup called Ripple Match. And around the time of the release of ChadGBT, I really was shaken awake quite abruptly by the rate of progress in AI. I led our initiatives at Ripple Match, integrating generated AI into the product.

101 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000665701263