**Nathan Labenz** (0:00)
Hello, and welcome back to The Cognitive Revolution. Smart robots, it's safe to say, have the potential to change daily life as much and perhaps much more than AI chatbots and coding assistants. But I often find that people tend to forget about robotics when reckoning with AI's overall impact. That's understandable in as much as robots aren't yet widely available for people to experiment and play with directly, but it's still a major blind spot in many forecasts. And so today, I'm especially excited to share my conversation with returning guests, Keerthana Gopalakrishnan and Ted Xiao, researchers at Google DeepMind and two of many authors of the recent Gemini Robotics Technical Report, which describes Google's recent work to bring AI into the physical world. In our first conversation, two years ago now, Keerthana described robotics as being in its GPT-2 era. Now she puts it somewhere in the range of GPT-3 to 3.5. Qualitatively, that is a huge difference. GPT-2 wasn't useful for much of anything, whereas GPT-35 was sufficiently mind-blowing as to create the chat-GPT moment. But still it wasn't capable or reliable enough to do all that much high-value work. At least not without fine-tuning on specific narrow tasks. As you'll hear, today's robotics models are in a similar phase of development. Architectures are simplifying as foundation models become more capable. Out-of-the-box generalization is improving, both in terms of tasks and different robot form factors. And the demos are highlighting increasingly impressive perception and motor control, with examples of robots using food-serving tongs, closing Ziploc bags, and even folding origami. So, how did they do it? Starting with Gemini 2.0, which, much like our recent episode on Google's AI Doctor and AI Scientist work, strongly implies significant improvement coming soon. The team created two distilled models, which work together to control the physical robots. The Gemini Robotics Embodied Reasoning Model runs in the cloud. It's responsible for high-level understanding, and it updates plans every 250 milliseconds, while a smaller Vision Language Action Model runs in part on the device and outputs low-level motor commands at 50 cycles per second. Reliability still isn't where it needs to be for mass deployment, but fine-tuning on specific tasks helps quite a bit. In some cases, with as few as 100 example demonstrations.
In addition to the details of this work, we also discussed the nature of the relationship between robotics hardware and models in general, how datasets have scaled to date and how that's starting to change, what the failures look like and how tolerable they are, and whether humanoids or other form factors will be the first robots to break through and move the needle on economic output. While there's of course still a lot of work to be done and many open questions to answer, the bottom line for now from my perspective is that trends suggest that robotics is consistently 3-4 years behind the LLM wave. If that continues, we might expect the GPT-4 moment for robotics in just the next 1-2 years, and from there we might well see, as we recently have with AI Chatbots, a rapid proliferation of interactive, intelligent, generalist robots across society. As always, if you're finding value in the show, I'd appreciate it if you'd share it with friends, write us a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. Of course, we welcome your feedback as well, either via our website, cognitiverevolution.ai, or by DMing me on your favorite social network. A quick reminder also, I'll be speaking at Imagine AI Live, May 28th through 30th in Las Vegas, the Adapta Summit, August 12th and 13th in São Paulo, Brazil, and the Enterprise Tech Leadership Summit, September 23rd through 25th, again, in Las Vegas. If you'll be at any of those events, please send me a message and let's meet up in person. For now, I hope you enjoy this update from the frontier of Google's robotics research, with Keerthana Gopalakrishnan and Ted Xiao, authors of Gemini Robotics.
Keerthana Gopalakrishnan and Ted Xiao, welcome back to The Cognitive Revolution.
**Keerthana Gopalakrishnan** (4:09)
Yay.
**Ted Xiao** (4:10)
Thanks for having us.
**Nathan Labenz** (4:12)
My pleasure. A year is a long time in the AI game, and it's been about a year since our last conversation. Obviously, a lot of stuff has happened, right? We've seen multiple new robotics companies founded and launched. We've got humanoid robots, at least on my Twitter feed, walking around all over the place. There's some new foundation models. We've got new kind of exquisite looking hands. I thought I'd maybe just kick things off by inviting you each to just share a super high-level zoomed out perspective on what has changed over the last year in robotics. Where were we then? Where are we now?
95 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000708843260