**Erik Torenberg** (0:00)
Hello, and welcome back to The Cognitive Revolution. The Cognitive Revolution is brought to you in part by Granola. Just yesterday, I happened to see Ramp's monthly report on the fastest growing software vendors, and the number two company that is adding the most new customers right now is Granola.
**Tom McGrath** (0:17)
Why?
**Nathan Labenz** (0:18)
Aside from advertising on the Cognitive Revolution, I would chalk it up to an extremely smooth and easy to use product experience. If you're listening to this show, there is a good chance that you could, in theory, build AI workflows that capture audio, transcribe it, and use it in downstream prompts and workflows. But can your teammates?
**Erik Torenberg** (0:38)
That, I think, is where Granola really shines. By delivering a polished product experience that anyone can immediately install and understand, and by introducing AI capabilities in the form of recipes made by trusted thought leaders, Granola is making AI accessible to everyone. See the link in our show notes to try my Blindspot Finder recipe and explore all of the ways that Granola can make your raw meeting notes awesome. Not just for you, but for everyone on your team, regardless of their relationship with AI. Now, today, I'm speaking with Dan Balsam and Tom McGrath, CTO and Chief Scientist of Mechanistic Interpretability Startup, Goodfire, who in less than two years since founding the company have assembled an all-star research team, landed a first wave of blue-chip customers, including a couple that discovered Goodfire via Dan and Tom's first appearance on the show back in August 2024, published a remarkable series of results, and most recently announced a $150 million Series B Fundraise at a valuation of $1.25 billion.
Along with the fundraise, they've announced a new pillar in their research agenda.
**Nathan Labenz** (1:49)
Intentional Design.
**Erik Torenberg** (1:51)
A push to expand the scope of what interpretability science can do, by complementing the current paradigm of reverse engineering how trained models work, with a new approach focused on understanding and shaping the loss landscape to control what models learn during training, and ultimately how they generalize. We begin with a discussion of interpretability developments broadly, with Tom emphasizing the shift from techniques like sparse autoencoders that transform a network's messy internal representations to sparse vectors where each node represents a distinct concept to newer approaches that attempt to understand the intricate geometric structures that these concepts inhabit within the model's latent space. From there, we dive into their plans for intentional design, and their first proof of concept, a technique for reducing hallucinations that use as a probe trained to detect hallucinations, both to steer the model at runtime and as a source of reward signal for additional reinforcement learning training.
Such training setups are not without controversy. People worry, understandably, based on results like OpenAI's obfuscated reward hacking, that models will simply learn to fool their monitors rather than truly correcting their bad behaviors. But Dan and Tom meet this concern head-on, agreeing that paranoia is a way of life in alignment research, acknowledging that intentional design techniques are immature and probably should not be used on frontier models today, while also arguing, first, that the pace of AI capabilities' advances really requires us to explore any and all possible paths to understanding and control, and second, that the specific details of the techniques really do make all the difference. In this hallucination reduction work specifically, the key trick they found was to run the hallucination detection probe on a frozen copy of the model during training, so that the modified model would hopefully find it easier to learn not to hallucinate than to find a way to evade detection. More generally, Tom asserts that a key principle is to avoid fighting backpropagation. Because models are such high-dimensional beasts, gradient descent will inevitably find ways around any attempt to prevent the model from learning what the loss function directs it to learn.
**Nathan Labenz** (3:59)
Winning techniques, therefore, must find ways to shape the loss landscape, so that the model naturally wants to learn what we need it to learn.
**Erik Torenberg** (4:08)
In the final part of the conversation, we discuss some of Goodfire's many other recent papers, including their work with PrimaMente, which suggested a new research direction by revealing that a state-of-the-art model for predicting Alzheimer's diagnoses was basing its predictions on the length of cell-free DNA fragments. We also discuss a project that showed that it's possible not only to determine which model weights are used for memorizing facts and which are used for more general-purpose reasoning, but that you can actually improve model performance on at least some reasoning tasks by removing the memorization weights from the model entirely.
Along the way, we also touch on how Goodfire intends to balance its need for business growth with its public benefit mission as they decide what research to publish and when. Briefly consider how well we should expect today's interpretability techniques to work on new and different architectures, get Dan's thoughts on the possibility of AI consciousness, and much more. As usual when I catch up on interpretability, I left this conversation really impressed by how much progress has been made so quickly, but also really mindful of just how vast neural networks are and how much we still have left to discover and understand. With that, I want to thank Dan and Tom for giving me another chance to drink from the Goodfire Research Firehose, and I hope that you learn as much as I did from this survey of mechanistic interpretability advances and introduction to the new paradigm of intentional design with Dan Balsam and Tom McGrath of Goodfire. Dan Balsam and Tom McGrath, CTO and Chief Scientist at Goodfire. Welcome back to The Cognitive Revolution.
95 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000753356333