**Erik Torenberg** (0:00)
Hey everyone, Erik here.
At Turpentine, we're building the first media outlet for tech people by tech people. We produce the show you're listening to right now. We're also building a network product that is bringing together the best founders and executives in the game, from all stages.
The Turpentine Network is a high trust space where exceptional people can talk. We already have over 400 members and growing. The community benefits include tactical advice, tech stack recs, hiring referrals, in-person events, an investor database, and exclusive perks. If you want to join, you can apply at the link in the description.
**Nathan Labenz** (0:35)
Hello and welcome to The Cognitive Revolution, where we interview visionary researchers, entrepreneurs, and builders working on the frontier of artificial intelligence. Each week we'll explore their revolutionary ideas, and together we'll build a picture of how AI technology will transform work, life, and society in the coming years. I'm Nathan Labenz, joined by my co-host, Erik Torenberg. Hello and welcome back to The Cognitive Revolution. Today I am thrilled to share my conversation with Albert Gu, assistant professor at CMU, co-founder of Cartesia AI, and leader of the state space model revolution.
I'm always looking for something that could fundamentally change the AI landscape. And with that mindset, I was one of the very first to recognize the importance of the selective state space mechanism that Albert and his co-author Tri Dao introduced in their now famous Mamba paper last December. While a number of other projects had already delivered subquadratic scaling properties at the time, Mamba was the first architecture to also outperform transformers on key metrics. This result inspired my obsession with state space model research, which has so far produced my original Mamba emergency pod, a shockingly popular two-hour monologue in which I declared the end of the transformer era and the beginning of the new mixture of architectures era, and also a two-part Mamba Palooza episode back in March, which I created with Jason Moe, who chronicles state space model developments online at statespace.info, and who joins us again today. Both of those episodes are good background for today's conversation, and with that context, you'll understand why it was such an honor to get Albert on the show. We began with a discussion of where his ideas come from, in terms of intellectual history, practical motivation, theoretical inspiration, and ultimately a sort of black box design intuition. After that, we go deep into the technical details of Mamba and the more recently published Mamba 2, which, very notably and a bit surprisingly from my perspective, achieves better results with a simpler and somewhat less expressive core mechanism, which happens to allow for more efficient training on modern hardware. As you'll hear in that part of the conversation about all these various trade-offs, I finally cleared up one lingering but important misconception that I had about the Mamba architecture. I thought that it was GPU SRAM size that limited the size of the all-important internal state, when in fact it was the speed of computation of the original selective state space mechanism that was the actual practical constraint. This is a minor point, but if you hope to have frontier understanding, let alone insight, it is super critical to understand which resource constraint is the overall system bottleneck at any given point in time. And so I'm very glad to have cleared that up.
From there, we move on to briefly discuss the explosion of Mamba-inspired literature. Jason reports that there are now 267 papers and projects downstream of Mamba, and they've introduced a number of interesting innovations, including multiple approaches to scanning images and other non-sequential data, and also the use of multiple internal states. From there, we explore a couple of future directions for this research, including the opportunity to train models on super long context datasets, which Jason has been pushing forward independently in the background, and also the possibility of moving toward more expressive, but presumably also more expensive to train mechanisms. We even get to my favorite idea to speculate about, the potential to design multi-state systems where each state plays a specific role. Albert said that he is interested in pursuing this direction, as he thinks it could perhaps deliver both superior performance and easier interpretability.
We went over 90 minutes in this recording, and only briefly touched on the commercial applications of Albert's work at Cartesia at the very end. If you want to hear more about that, I can also recommend his recent appearance on the NoPriors podcast, which covers almost entirely different ground than we cover here today.
Finally, a brief reflection. For listeners who are relatively new to AI and still building taste and intuition, I hope this episode is encouraging. Just three years ago, I had a casual interest in AI, but really very little technical depth and no expectation that I could contribute to the broader conversation. Today, I'm proud to look back on my response to Mamba and feel like I got the big things really right. And also to reflect on this conversation and note that Albert agreed with some of my current intuitions on the path forward.
102 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000661124985