Scaling "Thinking": Gemini 2.5 Tech Lead Jack Rae on Reasoning, Long Context, & the Path to AGI artwork

Scaling "Thinking": Gemini 2.5 Tech Lead Jack Rae on Reasoning, Long Context, & the Path to AGI

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

April 5, 2025

In this illuminating episode of The Cognitive Revolution, host Nathan Labenz speaks with Jack Rae, principal research scientist at Google DeepMind and technical lead on Google's thinking and inference time scaling work. They explore the technical breakthroughs behind Google's Gemini 2.
Speakers: Nathan Labenz, Jack Rae
**Nathan Labenz** (0:00)
Hello, and welcome back to The Cognitive Revolution. Today, I've got the honor of speaking with Jack Rae, Principal Research Scientist at Google DeepMind and technical lead on Google's thinking and inference time scaling work. As one of the key contributors to Google's blockbuster Gemini 2.5 Pro release, Jack has tremendous insight into the technical drivers of large language model progress and a highly credible perspective on the path from here to AGI. Gemini 2.5 Pro, as I'm sure you know, marks a significant milestone on Google's AI journey. It's the first time that many observers, myself included, would rank a Google model as the number one top-performing model across many important dimensions. And this is not just about topping leaderboards. In my initial testing of Gemini 2.5, which I conducted before Google's PR team reached out to schedule this interview, I experienced one of those rare moments where a model significantly exceeded my expectations, forcing me to re-evaluate my sense of what's possible today and inviting me to re-imagine my workflows to take advantage of its unique strength in not just accepting but actually demonstrating incredibly deep command of hundreds of thousands of tokens of input context. This is a practical step up that I could feel almost immediately, so naturally I jumped at the chance to talk to Jack about all the work that went into it and how he understands the current state of play along a bunch of critical conceptual dimensions. We begin by asking why techniques like reinforcement learning from correctness signals appear to have suddenly started to work so effectively across the industry. Does this represent a proper breakthrough or is this more a culmination of steady incremental progress that has finally crossed important thresholds of practical utility? We also unpack the reasons that nearly all frontier model developers are releasing similar reasoning or thinking models in such a short period of time. Is this simultaneous invention driven by obvious next steps, or is there more cross-pollination somehow happening behind the scenes? We then consider the relationship between reasoning and agency. Will these reasoning advances translate to agenda capabilities, or is something more still needed? From there, we look at the role of human data in shaping model behavior. How does Google think about collecting human reasoning and step-by-step task processing data? And how intentional has Google been in training models to follow recognizable cognitive behaviors versus letting them develop their own problem-solving approaches during the training process? We also exchange intuitions about the relationship between models' internal feature representations and the patterns of behavior they use to leverage them. Consider whether reasoning in latent space should scare us or can be made safe via mechanistic interpretability, and discuss whether the application of reinforcement learning pressure to the chain of thought itself should be avoided, as OpenAI recently argued in their obfuscated reward hacking paper. Finally, we will discuss the roadmap from our current capabilities to AGI. What are the remaining bottlenecks? Do we need a memory breakthrough, or will continued scaling of context windows be enough to overcome all practical limitations? And should we expect deep integration of more and more modalities, as we've recently seen with text and image?
Throughout our conversation, Jack provides thoughtful, nuanced responses that absolutely should help us improve our understanding of today's AI systems, the work going on inside Frontier Labs, and the overall trajectory of AI development. Personally, I leave this conversation with the sense that for most developments we see from the Frontier Labs, the simple explanation is the best one. There's still a lot of low-hanging fruit left in large language model development. Researchers have internalized the bitter lesson and are trying to keep their approaches as simple and scalable as possible. And the rapid progress we observe is mostly the result of pursuing pretty obvious high-level conceptual directions and then methodically chipping away at the practical engineering challenges required to make them work at scale. The teams involved, as you'll hear, are seriously concerned with developing the technology safely, but are also feeling both a high level of genuine excitement and competitive pressure that keeps them moving forward as quickly as possible. As always, if you're finding value in the show, and I definitely think this is one of the higher alpha episodes we've done, we'd appreciate it if you'd share it with friends, write a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. And considering that the future is radically uncertain and the stakes are crazy high, with outcomes from a post-scarcity disease-free utopia, to an existential catastrophe, or even outright human extinction, all live possibilities in just the next two to twenty years, I take my responsibility in making this show extremely seriously, and I earnestly invite your feedback and suggestions. You can reach us either via our website, cognitiverevolution.ai, or by DMing me on your favorite social network. Now, I hope you enjoy this insider's perspective on scaling large language model thinking and the path from here to AGI with Jack Rae, Principal Research Scientist at Google DeepMind. Jack Rae, Principal Research Scientist at Google DeepMind and technical lead on Google's thinking and inference time scaling work. Welcome to The Cognitive Revolution.

63 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000702326730