**Sam Charrington** (0:04)
Hey, I'm Sam Charrington. Welcome to another episode of the TWiML AI Podcast.
**Swix** (0:08)
And I'm Swix. This is a special episode of the Latent Space pod with TWiML at Google IO. Welcome.
**Logan Kilpatrick** (0:16)
Thanks for being here. Thanks for hanging out with us. I'm excited.
**Swix** (0:19)
Logan, you are our first guest. You came back remotely a few months ago, and now you're back here. You're a lot of the face of the AI studio, basically, that a lot of people are using. I'm using it. And I think it's a really welcome change for people being more accessible with the rest of the Google suite. And Shrestha, you've been... I actually don't super know your role.
I just generally have you pegged as PM of the API team with a particular focus on live.
**Logan Kilpatrick** (0:46)
Shrestha runs the show behind the scenes. Behind the scenes is the most public face running the show behind the scenes. Model launches, the live API, generally all the stuff that's happening in the API stresses hard work.
**Shrestha Basu Mallick** (0:59)
Thank you for that, Logan. But I think everyone knows who really runs the show. There's public evidence there. But yeah, I work with Logan and a few other excellent PMs, but I lead the API side of the house.
**Swix** (1:14)
There's a lot of announcements. I think a lot of people have done their recaps. What are you guys' personal highlights over IO?
**Logan Kilpatrick** (1:20)
I'll break the rule and I'll give two that are not the sort of big, big flashy ones. I think the two that I think developers are going to be super excited about. One, thinking budgets coming to 2.5 Pro. And you'll be able to disable thinking as well. So if you just want 2.5 Pro as like a raw non-reasoning model, we'll have that hopefully in early June. And then thought summaries. So we've had this debate internally about like, do we need to show full thoughts? Do developers want full thoughts? I think developers say they want full thoughts. We have thought summaries right now as a sort of step in that direction. It'll be really interesting to find out and get the feedback around like, what are things that work with thought summaries? What are the things that don't work with thought summaries? I was reading some threads last night about like, thought summaries are now live in cursor as well. And people were sort of reacting to having summaries versus not full thoughts. So it'll be interesting to see, but I'm excited for both of those things. Because thought summaries are live now. Thinking budget for 2.5 Pro will land with the GA model in a couple of weeks.
**Shrestha Basu Mallick** (2:18)
Yeah, and I should say we already do have thinking budgets in 2.5 Flash. I do think, you know, with all of the features that we are releasing on top of our thinking models, summaries, budgets, I think this is our way of, you know, you have the models, but then we want to give developers as much control as they can on top of models. But coming back to your question about my favorite feature, it's really hard to pick because, like, all of these features we've been trying to push out for weeks. But I think native audio output is a first thing.
**Swix** (2:50)
I was just saying that with Quinn. Yeah.
**Shrestha Basu Mallick** (2:52)
Yeah, it's a personal highlight. I actually, Quinn and I have been playing with it together for a bit as well. I think, especially with all the, obviously the voices sound great. The fact that it can switch in and out of languages. So Matt Veloso, our boss, actually has a demo on Twitter where it actually speaks Klingon, even though that's not an officially supported language. But, you know, I speak Bengali. Just being able for it to switch into and out of Bengali and English, that's been special. And then if I get to pick another one, it's I'd say we released a new tool called URL context. And the idea is that you can use it by yourself or pair it with search to retrieve more in-depth information from webpages in a way that's respectful of our publisher ecosystem, of course. And I think this will unlock new use cases, like if people want to build their own version of a research agent, which is something developers ask us for a lot.
**Sam Charrington** (3:52)
It's worth mentioning that just prior to IO., there was a ton of new interesting new capability, including the update to Gemini 2.5 Pro, as well as the implicit context caching, which I know a lot of folks are waiting for.
21 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000748427981