Opening AI's Black Box with Prof. David Bau, Koyena Pal, and Eric Todd of Northeastern University artwork

Opening AI's Black Box with Prof. David Bau, Koyena Pal, and Eric Todd of Northeastern University

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

April 5, 2024

In this episode, we dive deep into the inner workings of large language models with Professor David Bau and grad students Koyena Pal and Eric Todd from Northeastern University.
Speakers: Erik Torenberg, David Bau, Koyena Pal, Eric Todd, Nathan Labenz
**Erik Torenberg** (0:01)
Hey, everyone, Erik here. We're really excited about a new AI show from Turpentine called Autopilot, hosted by Will Summerlin.
This podcast explores the adoption and rollout of AI in the industries that drive the economy, and the dynamic tech founders bringing rapid scalable change to slow moving industries, from law, to hardware, to aviation. Will interviews founders backed by Benchmark, Greylock, YC, and more to learn how they're automating at the frontiers in entrenched industries. Click on the link in the description to subscribe to Autopilot.

**David Bau** (0:31)
Machine learning worked okay, but didn't really work in profound ways until the last 10 years or so.
But now it's really working. It's really working remarkably well. So we're facing a new type of software that we cannot use traditional computer science tools and traditional computer science methods for dealing with how to ensure that it's correct, that it does the stuff that we want.

**Koyena Pal** (0:53)
Logit Lens is essentially the tool where it projects these intermediate states into the decoding layer. So essentially, we can just see for that current token prediction, what is the model currently thinking about? We wanted to start with a quote unquote simpler solution, like having some sort of linearity with decoding future tokens. That would have been like a nice final solution, but we realized that no, it's a lot more complex than that, at least at the moment.

**Eric Todd** (1:18)
We tested like 40 plus different tasks, and it seems like they're all mediated by this small set of heads, which is cool that it seems like the model has this sort of path that it communicates this task information. Even though in the prompts were never explicitly telling it what the relationship between the demonstration and the label are, it's able to figure that out and communicate it forward.

**Nathan Labenz** (1:41)
Hello, and welcome to The Cognitive Revolution, where we interview visionary researchers, entrepreneurs, and builders working on the frontier of artificial intelligence.
Each week, we'll explore their revolutionary ideas, and together, we'll build a picture of how AI technology will transform work, life, and society in the coming years. I'm Nathan Labenz, joined by my co-host, Erik Torenberg. Hello, and welcome back to The Cognitive Revolution. Today, I'm excited to share a conversation with Professor David Bau of Northeastern University and grad students, Koyena Pal, author of the paper Future Lens, and Erik Todd, who made a surprise appearance to discuss their recent paper describing function vectors. Professor Bau left Google to establish this group with the goal of cracking open the black boxes that are modern AI models, so as to understand the internal mechanisms that underlie their key capabilities and ultimately gain robust control of increasingly powerful systems. And indeed, starting originally with vision models and now working more on LLMs, the Bau Lab has done pioneering work in this area, delivering a remarkable series of mechanistic interpretability papers over the last couple of years. In this conversation, we go deep on a couple of the group's recent publications. First, Koyena walks us through her future lens work, which develops techniques that show that even mid-sized language models seem to be thinking multiple tokens ahead. Then, Erik Todd joins to tell us about function vectors, a complicated pattern of activity spread across multiple attention heads in different layers of the transformer that seem to enable in-context learning by encoding the nature of the task inputs and outputs. This work is fascinating for anyone who wants to develop their intuition for how today's models actually function under the hood. And they're a notable step on the path to overall AI interpretability, identifying new abstractions that translate low-level patterns of computation to higher-order behaviors. I was a fan of this work coming into the conversation, but I left as a big fan of Professor Bau as a research advisor as well. Throughout the conversation, he does a great job of connecting the dots between his students' individual projects, motivating the lab's broader interpretability agenda, and emphasizing key open questions in the field. As always, if you find this work valuable, please do share the episode with your friends. I'd especially suggest this one for anyone who's interested in mechanistic interpretability and considering a PhD in machine learning.
And please, don't hesitate to share your feedback or suggestions via our website, cognitiverevolution.ai, or by messaging me on your favorite social network.
Now, please enjoy this illuminating discussion about the inner workings of large language models with mechanistic interpretability researchers, Koyena Pal, Erik Todd, and Professor David Bau. Koyena Pal and David Bau from Northeastern University, authors of Future Lens. Welcome to The Cognitive Revolution.

**Koyena Pal** (4:32)
Thank you.

**David Bau** (4:33)
Really happy to be here, Nathan.

74 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000651514696