**Jerry Tworek** (0:00)
If I play football, for example, it looks very, very close to reinforcement learning. I kick a ball a lot of times, and every time I adjust it a little bit, and I see if it roughly matches what I wanted, and either some self-reinforcement had money. When I learn mathematics, it's very different type of thinking. It's like reading about hard concepts and thinking about them very deeply inside my head until things click, until I have them connected. And both of those in some way are learning from experience.
They are just very different. We probably are spending the most computer lab than ever on learning from experience, but reinforcement learning is not the end of learning from experience. And there will be better approaches that researchers will be coming up in the coming years on how to use that data.
**Sonya Huang** (1:00)
Jerry, Rohan, thank you so much for joining us today. The two of you are the founders of Core Automation, one of the hottest Neo labs in San Francisco right now. And before studying Core Automation, you led some of the most important research projects of the AI era. Jerry, you were VP at OpenAI, where you worked amongst other things on running the Strawberry and Reasoning teams.
And Rohan, you were one of two of the four pre-training leads at Gemini, and before that led a lot of the fundamental AI research at Google Brain, and were the fix-it guy across Google and then at Anthropic. And so between the two of you, you've seen more than your fair share of what the world looks like in terms of doing frontier research. And so I'm very, very excited to dig in.
Let's start with you, Jerry. You tweeted a very spicy take recently. The first step to replacing Transformers is appreciating deeply how far they were able to carry us. Is that a eulogy for the Transformer? What does that mean?
**Jerry Tworek** (1:57)
Thank you very much for inviting us here, Sonya. I feel like a lot of my interviews these days is explaining my tweets and what did I mean.
But appreciating Transformer means like understanding what it does well. You are not solving the problems that it is solving well. You have to focus on its weaknesses. You have to understand good parts and bad parts. And it's very easy in a lot of the work, what people are doing in architectures is trying to make Transformers cheaper and trying to make Transformer more efficient. I very rarely see people thinking about how do we make Transformers more powerful, trying to do more expressive. But seeing someone's weak parts and seeing someone's strong parts are almost the same thing. It's just understanding the shape of Transformer a little bit more. Well, I think right now we are in this stage. We got really, really good at training really, really big models. We mastered two algorithms. We mastered pre-training at a large scale. And we mastered reinforcement learning at a large scale.
And I'm asking myself a lot what is next in machine learning. I think at this moment what the bottleneck is to better models and to smarter systems is the architecture itself. It is this moment to revisit the train we've been riding for the last six years of trying to add more and more parameters to essentially two of the same operations, which is MOE and attention. And when I'm thinking about it, like where we are today and what we are doing, I am thinking a lot about what Codex and what Cloud Code are doing for us. And I am really, really appreciative of those systems and of the coding and of the workflow automation and of the systems, of the products that we have today, that we essentially have built over those six years of scaling. And I think this is the first step of thinking, what is the, if we want to work on the replacement, we need to like see where we are, what problems we have solved to like start seeing what the next stage is, what problems we haven't solved yet. What kind of are we missing? And this is kind of whenever I use Codex and I am successful at a task, I also start thinking, why didn't I try to push that thing harder? Whenever I come to work, there are a lot of things I do with Codex, but I still come to work. I still ask it to do certain things for me. And I'm always asking myself, why am I even needed there? Why is Core Automation its name and its concept is we want to be automating tasks? And why are those things not yet automated? Why is not Codex doing everything for me? And this is the question of the research, where we want to go. And with that research, I'm trying to think, what kind of models, what kind of systems do we need? What kind of qualities do we need?
38 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000778868298