Breaking: Gemini's Major Update - Search, JSON & Code Features Revealed by Google PMs artwork

Breaking: Gemini's Major Update - Search, JSON & Code Features Revealed by Google PMs

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

October 31, 2024

Nathan interviews Google product managers Shrestha Basu Mallick and Logan Kilpatrick about the Gemini API and AI Studio. They discuss Google's new grounding feature, allowing Gemini models to access real-time web information via Google search.
Speakers: Nathan Labenz, Shrestha Basu Mallick, Logan Kilpatrick
**Nathan Labenz** (0:01)
Hey, everyone. Erik here. We've got something exciting in the works, and we want you to be the first to know about it. Turpentine, the network behind the show you're listening to right now, is launching a publication, and we're offering early access to our listeners. We'll have our biggest hosts and expert guests writing pieces and leverage our group chats for content inspiration. For an early preview, drop your email at the link in the show notes. You can also head to turpentine.co.
Now, on to the show. Hello, and welcome to The Cognitive Revolution, where we interview visionary researchers, entrepreneurs, and builders working on the frontier of artificial intelligence. Each week, we'll explore their revolutionary ideas, and together we'll build a picture of how AI technology will transform work, life, and society in the coming years. I'm Nathan Labenz, joined by my co-host, Erik Torenberg. As a developer, the journey from concept to production-ready large language model apps is fraught with challenges. Dealing with unpredictable language model outputs, hallucinations, and ballooning API costs can all be blockers to shipping your next AI-powered feature. That's where Advanced RAG comes in. With the new RAG++ course from Weights and Biases, you can overcome these hurdles and build reliable production-ready RAG applications. Go beyond proof of concept and learn how to evaluate systematically. Use hybrid search correctly and give your RAG system access to tool calling. Based on 21 months of running a customer support bot in production, industry experts at Weights and Biases, Cohere, and Weaviate show you how to get to a deployment-grade RAG application. This offer includes free credits from Cohere to get you started. Make real progress on your large language model development and visit wnb.me.cr to get started with their RAG++ course today. That's wnb.me.cr to get started with their RAG++ course today.
Hello, and welcome back to The Cognitive Revolution. Today, my guests are Shrestha Basu Mallick and Logan Kilpatrick, product managers for Google's Gemini API and AI Studio products. The occasion for our conversation is Google's release of a new grounding feature that allows Gemini models to tap into Google search results to get up-to-date information from the web at runtime. This comes just a couple of days after Google's latest earnings call, in which CEO Sundar Pichai noted that Gemini API usage has grown by a factor of 14 over just the last six months. To prepare for this conversation, I spent a couple of hours integrating Gemini into an application that I'm developing, and I found myself pondering a few big-picture questions. First, considering that various leaderboards consistently have Gemini at or very near the top, why does it still seem to be the third priority for most developers behind OpenAI and Anthropic? One hypothesis is that Google's platform is simply too complicated, or that the Gemini API is missing key features. But in my hands-on work, I did not find that to be the case. Using the Vercell AI SDK, I was able to integrate Gemini as a peer to OpenAI and Anthropic models quite quickly and easily. Second, should we expect large language model capabilities to converge or to diverge across providers? Considering this question through the lens of the Platonic Representation Hypothesis and taking seriously that large language models are learning ever more sophisticated world models and patterns of reasoning, it would seem that they are destined to converge. And yet, at the same time, OpenAI now has unique reasoning models, Claude is alone in its ability to use computers, and Gemini still has by far the longest context windows and now adds to that a native Google search integration. Finally, how and how much are teams within these companies thinking about competition? We know that big tech leadership is taking a game-theoretic approach, justifying tens of billions of dollars of investments on the ground that that's simply what it takes to make sure they're not left behind in the AI era. But what about at a product development level? Are they constantly thinking about the competition, or are they just trying to build the best products that they can? Logan and Shrestha had thoughtful responses to these questions and more as we got into a number of details on the Gemini API and AI Studio, including code execution, prompt caching, the incredible value of Gemini Flash, and the generous free tier that Google offers all in just an hour's time. As always, if you're finding value in the show, we'd appreciate an online shout out, a review on Apple Podcasts or Spotify, or just a comment on YouTube. We always welcome your feedback either via our website, cognitiverevolution.ai, or by DMing me on your favorite social network. For now, I hope you enjoy this real-time update on the Gemini API's latest features and the inside look at Google's AI API product strategy with Shrestha Basu Mallick and Logan Kilpatrick. Shrestha Basu Mallick and Logan Kilpatrick from Google, the Gemini API and the AI Studio team. Welcome to The Cognitive Revolution.

48 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000675237516