**Nathan Labenz** (0:00)
Hello, and welcome back to The Cognitive Revolution. Today, I'm excited to share my conversation with Jesse Hoogland and Daniel Murfet, founders of Timaeus, an AI safety and alignment research nonprofit that's pursuing an ambitious, mathematically rigorous, and fascinating approach to understanding the development and function of neural networks. Named after one of Plato's dialogues, Timaeus' work is based on Singular Learning Theory, or SLT, which applies algebraic geometry to statistical learning theory. Obviously, that's a mouthful, but the core premise of SLT is pretty intuitive. Training data determines the geometry of the loss landscape, which in turn determines which algorithms models learn in training, and ultimately how their behavior will generalize once training is done and they're put into actual use. The driving insight of SLT is that the super high-dimensional loss landscapes in which modern neural networks are optimized are not actually well represented by the smooth bottom valley-shaped surfaces that we often see depicted in figures. On the contrary, Daniel, who recently left a tenured professorship in algebraic geometry to pursue this work at Timaeus full-time, calls these representations maximally misleading and explains that in reality, loss landscapes are highly complex, jagged surfaces full of singularities, also known as degeneracies, which are directions in weight space that a model can move without changing its external behavior or loss core, but which nevertheless sometimes do involve a change to the model's internal circuitry such that the model might behave very differently in novel situations. This of course has profound implications for big-picture AI safety questions. To frame it in terms that would be familiar to Eliezer Yudkowsky readers from 15 plus years ago, the difference between a model that acts nice and friendly because it is fundamentally aligned to human values, and a model that acts the same way because it's learned how to please humans while actually pursuing its own goals, could be the difference between a super-intelligence-powered utopia and human extinction. And yet today, even as AIs become more and more powerful, we don't have reliable ways to tell the difference. Anthropic and good-fire style mechanistic interpretability has of course made great progress toward identifying the concepts that trained neural networks represent, and these days also offering some visibility into the circuits they use. But there's a very long way to go, and I definitely believe that there's plenty of opportunity for complementary approaches to strengthen our overall understanding. The Timaeus approach, which they call developmental interpretability, aims to understand how neural networks evolve through the training process, using a measure called the Local Learning Coefficient to help identify what might otherwise be invisible internal phase changes that could profoundly affect downstream model behavior. This line of work, like all approaches to understanding neural networks, is still pretty early in its own developmental history. But critically, the Timaeus team has shown that it can scale beyond toy models. Their latest work applies these techniques to 7-billion-parameter models and is able to identify critical phase change moments that correspond to the appearance of important functional circuits. So we might say that it's roughly at the toward monosemanticity stage and hope that with more engineering and compute, it will continue to scale to frontier models. If successful, this could perhaps prevent an episode like the one that happened in Cloud 4 training, where a certain safety dataset related to harmful system prompts was mistakenly left out of the data mix, causing the model to generalize in such a way that it followed rather than refused harmful system prompts. Anthropic caught that problem with behavioral testing and patched it, but the hope for developmental interpretability is that such things might be caught by instrumentation during the training process before they ever seriously affect model behavior. In the best case scenario, this could help us move beyond the trial and error phase of neural network training and toward a more principled engineering like discipline, where specific datasets are used at specific times for specific purposes, leading to predictable and reliable results. Now, this is all high-concept, mathematically sophisticated work, and this conversation, to be honest, is really just an introduction. I did my best to take my time to develop my own intuitions for what's going on inside a neural network during training, to compare and contrast some of the phenomenon that Jesse and Daniel described to other things like Grokking that we've previously covered, and to grapple with questions of how much generalization we really want from neural networks, and how in some cases too much generalization could be very harmful. I found this conversation fascinating throughout, and while it will stretch your brain in a bit of a different way than most of our episodes, I expect that for some of you it will immediately be among your favorite episodes that we've ever done. Now, I hope you enjoy this introduction to developmental interpretability, a new approach to understanding neural networks, with pioneers Jesse Hoogland and Daniel Murfet, founders of Timaeus.
79 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000713591368