Topics: Technology
**Nathaniel Whittemore** (0:00)
What if I told you that figuring out what parts of your work you should be getting AI to automate was a simple math equation? This week, two new products came online that make getting AI to do work for you much simpler. GrokBot gives users the ability to teach it a task by manually recording them doing something, while ChatGPT's computer history watches how you work and learns over time. Together, these represent the shift of the biggest challenge in AI moving from capability to context. But as these new features come online, you still have to figure out which part of the work you want AI to automate. The work best suited for AI deputization is frequent, time-consuming, teachable, easily verifiable, and doesn't require you to have been the one to do it to be successful. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI.
All right, friends, quick announcements before we dive in. First of all, thank you to today's sponsors, KPMG, Blitzy, Harbor, and HyperAgent. To get an ad-free version of the show, go to patreon.com/aidailybrief, or you can subscribe on Apple Podcasts. To learn more about sponsoring the show, send us a note at sponsors at aidailybrief.ai. And one more thing before we get into the headlines, including a new model from Gemini, a lot of the shows this week, including today's show, have to do with the newly launched GrokBot. And for those of you who are doing our AI Summer Adventure, choose your own adventure learning program, we've just posted a new pop-up destination, IE Project, all about testing out GrokBot. You can sign up for free at summeradventure.ai and walk through how to use this always-on teammate that I think is the simplest, cleanest version of OpenClaw we've ever had. Again, you can find that at summeradventure.ai, but now let's get into the headlines.
It's almost like Google heard us yesterday talking about how many people were talking about SpaceX AI as though it had completely usurped Google in the pantheon of serious frontier model companies. On Thursday, Google released Gemini 3.7 Flash. And while it is neither the much-delayed Gemini 3.5 Pro nor the now increasingly anticipated Gemini 4, it does play in a different category of efficiency that's becoming a higher and higher consideration, especially for serious and advanced users.
So let's start with what's good about this model. It appears to be very, very fast. During testing from Artificial Analysis, the model ran at 340 tokens per second, which is an entirely different category than anything else. It's more than twice as fast as GPT-56 Luna and even a bit faster than NVIDIA's new Nemotron 3.5 Lightning. On the benchmarks, Google made some solid gains over 3.6 Flash, most notably improving their score on coding benchmark DeepSweep from 48.6 to 65.3%.
And with this model, Google is also slashing prices by half, making it a little more cost-effective than its predecessor. Unfortunately, just like 3.6 Flash, by optimizing for speed, the model kind of ends up in a strange no-man's land. Even with the cost reduction, the model still costs $0.40 per task on the artificial analysis benchmark run. That makes it the same price as MuSpark 1.2, slightly more expensive than models like Nemotron 3 Ultra or GLM 5.2, and around eight times more expensive than the ultra-cheap models like GPT-56 Luna.
At the same time, it feels like the model just isn't strong enough on the benchmarks to justify the cost difference. Cognition pointed out that 3.7 Flash has Sonnet 5 coding performance for less than half the cost, but to put it mildly, Sonnet 5 has not been a hit. Most users are either paying a little more for a frontier model or looking for a much cheaper model. None of which is to say that Gemini 3.7 Flash is bad. It just sits in an uncomfortable middle ground right now on the cost-per-intelligence trade-off. And yet, I think it would be a mistake to assume yet that we really understand how on an aggregate these behaviors are going to shape out. In practice, some users are reporting increased utility, particularly with the boost in coding. Brandon Galang of Versel wrote, Ignore the FUD on Gemini 3.7 Flash. It actually sits on the Pareto frontier. It still technically gets edged out by 5.6 Luna, but that's extremely deceptive. 5.6 Luna is a much smaller model, and while my team has cut over a lot of production workflows to it, I personally would not turn to it for coding tasks. Gemini models have always had strong pros in multimodal understanding. I'm actually feeling somewhat iffy on Grok 4.6 from a model behavior standpoint, so for now I'm giving Gemini 3.7 Flash a try as my daily driver and execution model. Analyst Max Weinbach agreed, saying, I'm really enjoying Gemini 3.7 Flash and antigravity. It's really fast, it seems to be really good, and the usage limits are insanely high.
26 more minutes of transcript below
Thousands of transcripts fetched by people building searchable podcast archives
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/YOUR_EPISODE_ID