**Erik Torenberg** (0:01)
Hey, everyone, Eric here. We're really excited about a new AI show from Turpentine called Autopilot, hosted by Will Summerlin.
This podcast explores the adoption and rollout of AI in the industries that drive the economy, and the dynamic tech founders bringing rapid scalable change to slow moving industries. From law, to hardware, to aviation, we'll interviews founders backed by Benchmark, Greylock, YC, and more to learn how they're automating at the frontiers and entrenched industries. Click on the link in the description to subscribe to Autopilot.
**Nathan Labenz** (0:32)
Hello, and welcome back to The Cognitive Revolution. Today, we're doing something a bit different, which I am really excited about, and hope will become a regular part of the show.
With everything in AI going exponential all at once, I've been challenging myself to find new ways to keep my AI worldview as accurate and up-to-date as possible, and also to better help all of you to stay on top of the most important trends.
I really love doing the deep dive interviews with researchers and entrepreneurs and absolutely will continue to do them. But I increasingly feel that what is scarcest and therefore most valuable in today's world is a zoomed out perspective that attempts to make sense of whole research subfields and new emerging market sectors. And so today, that's exactly what we're going to try to do. My co-pilot on this adventure is Jason Meaux, fellow AI scout and creator of the website statespace.info, where he tracks Mamba and other statespace model research. Together in this two-part episode, we'll cover the first 30 Mamba-Inspired research publications, which Jason identified in just the first 90 days after the original Mamba paper. If you haven't heard my original Mamba monologue episode from December, I would recommend listening to that one first. As that episode presents a high-level description of how transformers function, what capabilities they are missing, and why I believe the new selective state-space mechanism introduced in the Mamba paper marks the beginning of what I'm calling the Mixture of Architectures era. All that is really important background information for the developments that we'll be discussing today. And indeed, there have been many interesting developments. The work we'll cover today explores the mechanistic function of the Mamba architecture, the relative strengths and weaknesses of the selective state-space and attention mechanisms, application of Mixture of Experts strategies to the Mamba architecture, use of Mamba models for image segmentation and other computer vision tasks, attempts to realize the promise of much longer context windows, and lots more. Along the way, we identify a number of important themes, including the use of multiple internal states, which was one of my big predictions from the December episode, the need to cast all input data types as sequences, and the use of multiple different scans over the input data to accomplish this, the emerging dominance of hybrid architectures, another big prediction from December, and the many open questions around how best to handle and what more we might be able to do with those all important hidden states.
This was a fun and intellectually invigorating conversation, and I am truly grateful to Jason for all his hard work collecting the research and joining me to break it down. He even contributed to the editing process. This was definitely going above and beyond, but he has created a clearer, more information-dense listening experience for you as a result. As part of that, he even added key figures to the video version of this episode. If you want to see those as you watch, and I do think it can be quite helpful at times, please visit our YouTube channel.
We packed as many key topics as we could into these two hours and change, but even so, we were not able to cover everything. And on top of that, in just the short time since we recorded, there have already been a bunch more papers, including several that seem quite important. With that in mind, we see this as just the start of an ongoing series, and we look forward to bringing you another update on this dynamic area in the near future. For now, as always, we invite your feedback on the show. If you enjoyed this format of surveying a whole research literature, please do share it online. If there are other fast-moving AI research areas you'd like to see us cover in a similar way, we welcome your suggestions. And if you personally are in position to do a similar project, I would love to work with you on it. I've got projects in progress exploring the scale of data that will be required, as well as the scale of data that's available to produce next-generation models, the application of generative AI to biology, and the latest developments in brain-computer interfaces. And really, I would love to do so many more like this.
63 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000650899161