**Erik Torenberg** (0:00)
Turpentine is a network of podcasts, newsletters, and more covering tech, business, and culture, all from the perspective of industry insiders and experts.
We're the network behind the show you're listening to right now.
At Turpentine, we're building the first media outlet for tech people by tech people. We have a slate of hit shows across a range of topics and industries, from AI with Cognitive Revolution, to Econ 102 with Noah Smith. Our other shows drive the conversation in tech with the most interesting thinkers, founders, and investors like Moment of Zen and my show Upstream. We're looking for industry leading hosts and shows along with sponsors. If you think that might be you or your company, email me at erikaturpentine.co. That's E-R-I-K at turpentine.co.
**Nathan Labenz** (0:45)
No less than Imad Mostak from Stability said, brilliant researchers like this literally knock 10 percent off of global training compute needs. With these improvements, which are impossible to predict. 10 million tokens starts to give you the opportunity to put whole bodies of literature into a single token. I mean, the great Gatsby famously fits into Claude's 100K. Now you're talking perhaps about a 100 books with full attention considered for the next token generation. If this allows the models to make those connections at such huge length, this could be where you could start to see tipping into superhuman performance of learning things that experts don't know. Hello and welcome to The Cognitive Revolution, where we interview visionary researchers, entrepreneurs, and builders working on the frontier of artificial intelligence. Each week, we'll explore their revolutionary ideas, and together we'll build a picture of how AI technology will transform work, life, and society in the coming years. I'm Nathan Labenz, joined by my co-host, Erik Torenberg. So then the next one is Think Before You Speak, is the name of this paper, and Training Language Models with Paws Tokens. So Think Before You Speak, Training Language Models with Paws Tokens. This comes out of Google and again Carnegie Mellon Collab, and this was, I guess, a student from Carnegie Mellon who's interning at Google. It's amazing how many of these papers are like couple-month processes, and you can do this kind of stuff in the context of a summer internship these days.
Amazing. So what do they do here? They start off with an observation, which is a pretty simple one on some level. I'll quote this from the paper. Language models generate responses by producing a series of tokens in immediate succession. The k plus 1th token is an outcome of manipulating k hidden vectors per layer, one vector per preceding token.
What if instead we were to let the model manipulate, say, k plus 10 hidden vectors before it outputs the k plus 1th token? So basically just like, especially in the early going, but kind of anytime, you only have a certain amount of computational space to move information around. And maybe that's just not enough, or maybe more could be beneficial.
First thing that comes to mind for me when I hear something like that is, I think that's a lot of what's happening in the chain of thought type prompting. Certainly that's been hypothesized that by giving the model time to think, you give it time to kind of work its way through things and hopefully summon the right reasoning.
And then you get better results. So certainly empirically we see that you get better results. Well, this is now saying, okay, what if we just gave it extra space, but we didn't make it do anything with that extra space? I'm not asking for reasoning. I'm just giving it a pause token that it can just literally put a pause in when it feels like it needs to, and how exactly that gets decided is a bit of a black box.
And there's some trade-offs here, I think, for sure. But give it that opportunity to just take a pause when it needs to. Now it can just process information a little bit more.
It could potentially do a couple of pause tokens in a row if it needs to. They suggest, what about 10, k plus 10? So 10 extra vectors that it can kind of manipulate and move information back and forth between. Does that give us the opportunity for better performance? And obviously, this makes the research round up because indeed they show that they are able to improve performance on this.
So first thing on this, I was like, boy, I saw that one coming.
We covered a little bit of the backspace paper on a couple earlier episodes. And this kind of combines some concepts from the last one too. In the backspace one, when they got out of distribution, as measured by not having high confidence on any next token, then they started to train the model to use the backspace to go back and be like, well, we must have gotten off the rails here a little bit because now we're not confident of what to do next. So let's instead go back, try that one again, maybe make a different prediction this time, and then maybe that will lead us toward something where we can feel more confident. When I saw that, I was like, well, it seems like you could probably have a, if you can go back one, you could probably just add a padding one too, and just kind of quietly think to yourself. Sure enough, 60 days or so later, here's the publication. Was this inspired by that? I'm not sure. It was kind of on the border where there was just enough time for them to have done it in response to that, but I would guess, honestly, they probably had the idea before. So it's probably independent, kind of parallel lines of thinking.
47 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000631977765