Training Zamba: A Hybrid Model Master Class with Zyphra's Quentin Anthony artwork

Training Zamba: A Hybrid Model Master Class with Zyphra's Quentin Anthony

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

October 30, 2024

In this episode of The Cognitive Revolution, Nathan dives deep into the world of state space models with returning co-host Jason Meaux and special guest Quentin Anthony, Head of Model Training at Zyphra.
Speakers: Erik Torenberg, Jason Meaux, Quentin Anthony, Nathan Labenz
**Erik Torenberg** (0:01)
Hey, everyone. Eric here. We've got something exciting in the works, and we want you to be the first to know about it. Turpentine, the network behind the show you're listening to right now, is launching a publication, and we're offering early access to our listeners. We'll have our biggest hosts and expert guests writing pieces and leverage our group chats for content inspiration. For an early preview, drop your e-mail at the link in the show notes. You can also head to turpentine.co/exclusivedashaccess. Now, on to the show.

**Jason Meaux** (0:29)
The future of AGI will involve a combination of cloud and on-device deployment.

**Quentin Anthony** (0:36)
These large modelistic model companies doing like in Dr. KropenAI just can't really specialize to every single person on the planet. We think that you need to have your own set of weights, and changing a system prompt per person is not enough, right? We want to actually bake into the weights. You can make the model simulate learning faster than it really is by doing activation steering. If the user tells the model you're being too dry, then you can very quickly steer the activation to be a bit more fun. Until tonight, when you can bake into the model, I think it's got to be continual learning, and it's got to be per user. And the only way to do that is with weights on the phone.

**Nathan Labenz** (1:12)
Hello, and welcome to The Cognitive Revolution, where we interview visionary researchers, entrepreneurs, and builders working on the frontier of artificial intelligence. Each week, we'll explore their revolutionary ideas, and together we'll build a picture of how AI technology will transform work, life, and society in the coming years. I'm Nathan Labenz, joined by my co-host, Eric Thornburg. Hello, and welcome back to The Cognitive Revolution. Today, we're once again going down the state space model rabbit hole, with returning co-host Jason Meaux, who regular listeners will remember from our Mamba Palooza literature review and Albert Gu interview episodes, and Quentin Anthony, head of model training at Zyphra, a large language model startup that's just released their Zamba 2-7b model, which is built on a hybrid architecture that uses both the selective state space mechanism and the traditional attention mechanism, albeit with some notable tweaks relative to the standard implementation. In addition to sharing Zyphra's high-level vision for highly personalized on-device AI, Quentin was super generous with both his time and knowledge, sharing a wealth of practical lessons learned from the front lines of model training. Over the next two hours, we will cover the delicate architectural choices that balance efficiency and capability, the many practical challenges of training at scale, including choosing the right learning schedules for different phases of training, the nitty-gritty details of training hybrid architectures, including why Zamba models don't need positional embeddings, the Zamba models' use of shared attention blocks and internal LoRa adapters to maximize performance on the edge, the not-so-simple relationship between lost metrics and model quality and capabilities, as well as the challenges of context-linked extension, the Zyphra teams' experiments with different optimizers and why they're sticking with Atom for now, Quentin's intuitions about the relationship between model scale and lost landscapes, and finally, even their recent published work on tree attention, which offers important advantages over ring attention for multi-node training.
I have to say, I got a lot from this episode, and while it's technical enough that I wouldn't necessarily call it entertainment, I am confident that you will too. If so, we always appreciate it when folks take a moment to share the show with friends or write an online review on Apple Podcasts or Spotify, and we welcome your feedback via our website, cognitiverevolution.ai, or by DMing me on your favorite social network. Now, I hope you enjoy this highly technical conversation with co-host Jason Meaux and guest Quentin Anthony, Model Training Lead at Zyphra. Jason Meaux, returning guest and co-host and chronicler of Statespace Models at statespace.info, and Quentin Anthony, Model Training Lead at Zyphra, which has just released the new Zamba 7b SSM Hybrid Model. Welcome both of you to The Cognitive Revolution.

**Quentin Anthony** (3:58)
Thanks a ton. Great to be here.

**Nathan Labenz** (4:00)
Jason, regular listeners will know, as a partner in crime who's also obsessed with Statespace Models and the potential that they have to unlock new capabilities in terms of potentially long-term memory, extreme efficiency, all these kind of interesting things, and we've been down the rabbit hole together on that a couple times. So I was excited to have him back to help me with this conversation about the new Zamba model and all the ins and outs of that. So Jason, I'm going to ask you to lead the questioning today, and I'll be in the supporting role, but listening intently and probably jumping in with a few of my own follow-up questions along the way as well. How's that sound?

129 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000674971713