**Nathaniel Whittemore** (0:00)
Today on the AI Daily Brief, a security incident that has us asking, just how good is Gpt-6 really? Before that in the headlines, a new set of Google models, but not necessarily the ones that we wanted.
The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI.
All right, friends, welcome back to the AI Daily Brief Headlines edition, all the daily AI news you need in around five minutes. In all of the recent model talk, one lab that has been conspicuously absent is Google. It has now been months and months since we got any update from them on their Pro Series models, having to have contented ourselves with just smaller and faster models like 3.5 Flash. Yesterday's announcement did not bring 3.5 Pro, which has been rumored to be underperforming. Instead, we once again got a set of new variants of Gemini Flash. Tuesday's release was headlined by Gemini 3.6 Flash, and the big change is better token efficiency. On the artificial analysis benchmark run, the model used 17 percent fewer tokens than 3.5 Flash. Google also said that on some isolated benchmarks like DeepSui, they observed up to a 65 percent reduction in token usage. Now, this might be particularly relevant because one of the loudest complaints around the release of 3.5 Flash was that the model was significantly more expensive and heavy on token usage than its predecessors. Google appeared to have optimized for speed, but that left some people questioning exactly what the purpose of 3.5 Flash was relative to other models. And of course, with Chinese AI labs competing hard on cost efficiency, this left 3.5 Flash somewhat in no man's land, not good enough for high performance tasks and not cheap enough for low end tasks. Now, in addition to the reduction in token usage, some of the benchmarks suggest that 3.6 Flash has delivered a boost in performance. On coding tasks, it scored 49% on DeepSuite compared to 37% for 3.5 Flash, with similar levels of improvement observed across benchmarks for ML research, computer use and knowledge work. Then again, benchmarking from artificial analysis suggested that not all that much had changed. 3.6 Flash scored 50 on the Intelligence Index, which was the same score as 3.5 Flash. That said, AA did find a 50% speed boost and an 18% reduction in cost per task. Google is also cutting prices explicitly, reducing cost per million output tokens from $9 for 3.5 Flash to $750 for 3.6 Flash. Alongside 3.6 Flash, Google released 3.5 Flashlight and 3.5 Flash Cyber. Flashlight is the ultra-fast model designed for high-latency, agentic tasks, and compared to 3.1 Flashlight, the model delivered a 23-point jump on terminal bench 2.1 and almost doubled its score on GDPVal AA.
Now, none of these numbers are even close to frontier, but even before it became a thing in the wider enterprise world, Google had already started to compete for cost and efficiency-optimized types of models, which is clearly the game here as well. As the name suggests, Flash Cyber is a fine-tuned version of the model designed for cybersecurity work like bug hunting and patching. It scores 83.2% on the Cyber Gym benchmark, which actually puts it only a few points behind Mythos 5, GPT-56 Sol, and GPT-55 Cyber.
Now, presumably this model isn't quite as strong in other aspects of cyber work, but once again, having a cheaper and faster option for vulnerability mitigation could be a big deal. Flash Cyber won't actually see a general release, however, with Google making it available only to governments and trusted partners.
Now, of course, it's only been a short time, but people's first impressions of this model slate aren't great. Abacus AI's Bindu Reddy writes, Gemini 3.6 Flash scores below 3.5 Flash, so this seems worse than their last generation, more expensive than Grok and Luna. Very strange model release.
Leo at SynthWaved writes, Gemini 3.6 Flash benchmarks are out, and it's beaten by other models on code tasks and has only really consistently stated the art on vision and context benchmarks. But hey, 3.1 Pro is now so old that 3.6 Flash outperforms it across the board. Lassonde writes, they are so scared of training Pro. If Flash flops, they can at least say, we have a bigger model, this is not our best. Please lock in Google Bros. And indeed, the big question right now is what happened to Gemini 3.5 Pro? The model was anticipated all the way back to the IO conference in May, but at the time, CEO Sundar Pichai said that it was slated for a June release. June, of course, has come and gone with no 3.5 Pro, and there have been rumors of subpar performance pushing back the timeline. Google's Logan Kilpatrick insists it's still coming, posting on Tuesday, Gemini 3.5 Pro is currently testing with partners and we plan to make it broadly available as soon as it's ready. Perhaps a more exciting hint also came from Logan who added, we've started our most ambitious pre-training run yet for Gemini 4 and are excited by the progress.
26 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000777934844