**Nathan Labenz** (0:00)
Hello, and welcome back to The Cognitive Revolution. Today, after some 340 episodes, I am very excited to share the first episode that I've ever recorded in person, with fan favorite Logan Kilpatrick, member of technical staff at Google DeepMind, and Tulsee Doshi, Senior Director and Head of Product for Gemini Models. The occasion for this conversation is Google's annual IO event, where they're launching the new Gemini 3.5 Flash model, all sorts of agent infrastructure and AI product integrations, and plenty more. We recorded on Friday, May 15th, just a couple days before the event. While many at Google, including my brother Craig, who's giving a keynote on Wednesday, were working overtime to polish their demos and presentations, the overall vibe, at least compared to the rest of the AI space, was one of relatively relaxed confidence.
Why not? From 2024 to 25, Google grew annual revenue by $50 billion, as much as Anthropic is pulling in today. And they still have 25% of all global compute, the deepest pool of research talent anywhere, and the most comprehensive AI portfolio of any company, with top-tier positions not just in language models, but also self-driving cars, medical and life sciences, and robotics. So, after discussing the headline launches that they're announcing this week, which also include a new video generation model called Omni, which they hope will create a nano-banana moment for video, a new and improved and more agent-focused anti-gravity, and a product called Spark, which will bring more agentic functionality to the main consumer Gemini app, I really wanted to take a step back and dig in on Google's overall AI strategy and philosophy. We discussed their decision to lead with the Flash model, and more generally to emphasize the cost-adjusted performance Pareto frontier, whereas Anthropic and OpenAI are clearly much more focused on competing to have the single most capable model in absolute terms. We talk about how DeepMind is no longer shipping models in isolation and leaving it up to product teams to figure out how to use them, but instead now providing a robust agent harness, which should help elevate and standardize AI experiences across Google's vast product surface. We get into the weeds on questions like why context windows seem to have mostly stopped growing, why Gemini models' knowledge cut off is now more than a year ago, and whatever happened to that diffusion model line of work? Perhaps most importantly, we discuss how the team at Google relates to the AIs they're creating, how they're thinking about things like model psychology and welfare, and their views on recursive self-improvement, which, as you'll hear, is definitely a part of their plan, but not something that they seem to be so singularly focused on as other AI leaders.
Overall, I think this is a great window into the thinking that underlies Google's AI research and product development, which has clearly sustained the company's historic run far beyond the point that many analysts had written them off.
With that, I hope you enjoy my first ever in-person conversation with Logan Kilpatrick and Tulsee Doshi of Google DeepMind.
**Logan Kilpatrick** (3:16)
All right.
**Nathan Labenz** (3:17)
Well, we are here live at Google headquarters in the library at Gradient Canopy, the first ever in-person recording of the Cognitive Revolution. Logan Kilpatrick and Tulsee Doshi, welcome.
**Logan Kilpatrick** (3:27)
Thank you. This is an honor. I didn't realize this was the first in-person episode. Yeah.
**Nathan Labenz** (3:31)
350 plus. It's all been from my home office in Detroit until today.
**Logan Kilpatrick** (3:36)
That's awesome. Well, thank you for being here.
This is a crazy space, especially around IO. It's a zoo. Yeah.
**Nathan Labenz** (3:43)
It's always a good time here at Google HQ. So you may or may not remember the No Motes Memo. We've just passed the three-year anniversary. It was May 5, 2023
In the intervening three years, Google has added $3.5 trillion in market cap, which is more market cap than all but two other companies in the world, those two are NVIDIA and Apple. So the motes, I'd say, are holding up. Here we are at IO, and I'm sure there are going to be some exciting new things that will be deepening the motes. So first question, tell me, what are we launching this week to try to deepen those motes?
**Tulsee Doshi** (4:24)
A lot. So a lot of exciting stuff. So let's see, let's start with some of the modeling side of things, because that's really exciting. We have our 3.5 series coming out starting with 3.5 Flash at IO. We're really excited about 3.5 Flash, because I think Flash does this really awesome job of being at the sweet spot of being really smart, while also being really fast and really cost effective. And so Flash is incredible. It is like three times faster than other of the large models. It's significantly cheaper for being able to still drive these really awesome magentic encoding workflows, and we've been using it internally a lot, which has been really fun to kind of see that play out. So that's one big piece, which is 3.5 Flash, we're really excited about. We're also releasing Omni, Gemini Omni Flash, which is a video generation and editing experience. What's really exciting about Gemini Omni in general is it's our push towards being able to bring all modalities in and all modalities out. The first way this is really manifesting is in this video editing context. So you're going to be able to make really awesome videos. You're going to be able to put your own avatar into the videos, which is going to be awesome. I've been having a bunch of fun playing with that too.
52 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000768769266