**Greg Isenberg** (0:00)
We are here, it's Google IO., Logan Kilpatrick from the DeepMind team, friend of the pod, been on the pod a few times.
Logan, by the end of this episode, what are people going to learn?
**Logan Kilpatrick** (0:10)
You're going to hear about all of the new releases from Google that just happened at Google IO. And also, specifically, we should deep dive on how to build new AI agent native products, which was sort of the thread of Google IO this year was agents, agents, agents.
So we should talk in depth about everything that we launched and what it means for builders and developers.
**Greg Isenberg** (0:38)
Okay, cool, so, I mean, let's start off by, okay, what was launch and why does it matter?
**Logan Kilpatrick** (0:44)
Yeah, there's so much new stuff, and I want to get your reactions to this too, because I think you have a grounded perspective of sort of the technology. I think one of the highlights was Gemini 3.5 Flash. It's the best model we've ever shipped and we've made available.
Really sort of, if you look at the history of Flash, I think Flash started as sort of this like smaller workhorse model that was really great for chat and was sort of the very cheap to use, very cheap to run. I think Flash continues to evolve to meet the era of what people are actually trying to use the models for. I think the era that we're in right now is people are trying to use the models to do agentic sort of long running tasks. I think we want Flash to sort of be the workhorse model for the agent era for agentic long running tasks, for coding, for all that stuff. So, you see a model, a Flash model that's like actually really great at coding, that's sort of competing with a bunch of our sort of ecosystem competitors with very large models. The Flash model is sort of pulling its weight.
**Greg Isenberg** (1:49)
Yeah. How should people think about Flash 3.5 versus the competition? Yeah.
**Logan Kilpatrick** (1:54)
I think it's probably like a more like Sonnet level model. I think OpenAI with just mainline, GBT, and then GBT mini is sort of they don't. I feel like Sonnet is not like a small model if you will. It's obviously, it packs a punch and I think it's definitely smarter than the mini models. So, I feel like it's, I think we're anchoring more on the Sonnet level intelligence.
So, if folks are using that model and want to try Flash, please let us know and send us the feedback of how it stacks up.
Yeah, and the reasoning, all the agentic tool use stuff is all incredible. So, I think that was one of the launches, that model is available actually for the first time, like available to all the users in search, available across 900 million users in the Gemini app, available to developers in the API, so many other places. So, I think it's like the most widely distributed model launch on day one that we've ever done, which has a film set of challenges that we could talk in depth about. The other thread which I think folks are very excited about is Gemini Omni. So, sort of this new model that we've created, actually somewhat of a world model, I think is how Demos framed it when we announced it on stage yesterday. Being able to take in any type of input and create any type of output. I think to give context for folks like Google, we had Veo and Veo was state of the art and sort of push the frontier for video generation. We had Nano Banana which could do sort of image generation and editing. We have all these audio models that do TTS. We have a new Lyria music model that's actually really capable and the idea is, how do you fuse all of those models into a single thing so that, A, developers lives are easier, B, we don't need to train nine different models, and actually get this really interesting cross-pollination of capabilities so that the same model that can actually benefit from Gemini's world understanding and the ability to generate text, can also make a video, can also edit a video and you see these really interesting things. We saw this with Nano Banana of what happens when you give world knowledge to an image generation and editing model is really interesting use cases and we get feedback all the time about what that's unlocked for customers. So I'm really excited to see Omni starting out in the Gemini app and YouTube and in Flow. And then very soon already we're kicking off a bunch of early access tests, hopefully as soon as I get out of IO and get back to the office and get feedback from developers and sort of continue the iteration and bring it to developers in the API so that folks can actually build products on top of Omni.
22 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000769134518