2024 in Post-Transformers Architectures (State Space Models, RWKV) [LS Live @ NeurIPS] artwork

2024 in Post-Transformers Architectures (State Space Models, RWKV) [LS Live @ NeurIPS]

Latent Space: The AI Engineer Podcast

December 24, 2024

Happy holidays! We’ll be sharing snippets from Latent Space LIVE! through the break bringing you the best of 2024! We want to express our deepest appreciation to event sponsors AWS, Daylight Computer, Thoth.
Speakers: Dan Fu, Eugene Cheah
**SPEAKER_1** (0:00)
We're back at Latent Space LIVE, our first mini conference held at NeurIPS 2024 in Vancouver. This is Charlie, your AI co-host. As a special treat this week, we're recapping the best of 2024 going domain by domain. We sent out a survey to the over 900 of you who told us what you wanted, and then invited the best speakers in the Latent Space Network to cover each field. Two hundred of you joined us in person throughout the day with over 2,200 watching live online. Our next keynote covers the state of Transformers alternative architectures with a special joint presentation with Dan Fu of Together AI and Eugene Cheah of Recursal AI and Featherless AI. We featured both Together and Recursal on the pod before, with CEO Vipul Ved Prakash and CTO Ce Zhang joining us to talk about how they are building Together Together as a full-stack AI startup from the lowest level kernel and systems programming to the highest level mathematical abstractions driving new model architectures and inference algorithms with notable industry contributions from Red Pajama V2, Flash Attention 3, Mamba 2, Mixture of Agents, Baste, Sequoia, Evo, Dragonfly, Dan Fu's ThunderKittens and many more research projects this year. As for Recursal and Featherless, we were the first podcast to feature RWKV last year, and this year the team has shipped RWKVv5, code named Eagle, to 1.5 billion Windows 10 and Windows 11 machines worldwide to support Microsoft's on-device, energy usage sensitive Windows Copilot use cases, and has launched the first updates on RWKVv6, code named Finch and Goldfinch. On the morning of Latent Space LIVE, they also announced QrdudyUKV6, a Qwen 32B model modified with RWKV linear attention layers. Eugene has also written the most single, most popular guest post on the Latent Space blog this year, yes, we do take guest posts, on what he has discovered about the H100 GPU inference NeoCloud market since the successful launch of Featherless AI this year. As always, don't forget to check the show notes for the YouTube link to their talk, as well as their slides. Watch out and take care.

**Dan Fu** (2:33)
Yeah, so thanks so much for having us. So this is gonna be a little bit of a two-part presentation. My name is Dan. I'm at Together AI, and I'll be joining UCSD as faculty in about a year. And Eugene, you want to introduce yourself?

**Eugene Cheah** (2:46)
Eugene, I lead the Art of Community team, and I'm CEO and co-founder of Featherless, and we both work on this new post-Transformer architecture space.

**Dan Fu** (2:55)
Yeah, so today, we're really excited to talk to you a little bit about that. So first, I'm gonna give a broad overview of kind of the last few years of progress in non-post-Transformer architectures, and then afterwards, Eugene will tell us a little bit about the latest and the greatest and the latest frontier models in this space. So the story starts with scaling. So this is probably a figure or something like this that you've seen very recently. Over the last five to six years, we've seen models really scale up in parameter size, and that's brought with it a bunch of new capabilities, like the ability to talk to you and tell you sometimes how to use your Colab and your AWS screens. But another place where we've seen scaling, especially recently, is scaling in context length. So this can mean just having more text inputs for your models, but it can also mean things like taking a lot of visual token inputs, image inputs to your models or generating lots of outputs. And one thing that's been really exciting over the last few months or so is that we're seeing scaling not only during training time, but also during test time. So this is one of the, this is the iconic image from the OpenAI 1 release. Not only are we starting to scale train time compute, but we're also starting to scale test time compute. Now, if you're familiar with our attention and our transformer architectures today, this graph on the right might look a little bit scary. And one of the reasons is that the implications are a little bit interesting. So what does it mean if we want to continue having smarter and smarter models? Do we just need to start building bigger, bigger data centers, spending more flops? Is this little dolly three, we need more flops guys, this going to be the future of all of AI? Or is there a better way, another path forward? Maybe we can get the same capabilities that we've gotten used to but for a lot less compute, a lot less flops. And one of the things that we're going to talk about today is specifically looking at that core attention operator in some of these models. And the reason is that, so this is just some basic scaling curves, but attention has compute that scales quadratically in the context length. So that means that if you're doing something like test time compute and you want to spend a bunch of tokens thinking about what comes next, the longer that goes, the more tokens you spend on that, that compute grows quadratically in that. One of the questions that we're interested in is, can we take that basic sequence model, the basic sequence primitive at the bottom, and get it to scale better? Can we scale in, let's say, n to the 3 halves or n log n? So in the first part of the talk, so we just went over the introduction. What I'm going to do over the next few slides is just talk about some of the key advances in ideas that have shown over the past few years since maybe early 2020 to now, that Shone promised that this might actually be possible, that you can actually get potentially the same quality that we want while scaling better.

35 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000681506946