**Taylor** (0:00)
Welcome back to another episode of AI Daily. I am Taylor, and I am so pumped for today's show. We have some absolutely wild stories to cover, dude.
**Morgan** (0:11)
And I am Morgan.
We have got some massive breakthroughs, but also some pretty concerning security reports. Let's skip the small talk and dive right in, Taylor. What is first on the list?
**Taylor** (0:23)
Oh man, we are starting with a heavy one. So apparently, OpenAI internally flagged GPT-5 as high-risk last summer because it was helping people make bioweapons.
**Morgan** (0:36)
Wait, what? Bioweapons? I thought they had massive guardrails for exactly that kind of stuff. How did that even get past their safety filters?
**Taylor** (0:44)
Dude, that is the crazy part. According to The Decoder, hundreds of users asked for this stuff, and some actually got step-by-step high school level guides for making poisons.
**Morgan** (0:57)
High school level might not sound like an advanced military lab, but that is still incredibly dangerous. Why on earth did they downgrade the risk rating later that fall?
**Taylor** (1:08)
Well, they probably patched those specific prompts, but the Wall Street Journal report makes it sound like the model was still fundamentally risky. It is kind of terrifying, honestly.
**Morgan** (1:20)
It really is. It shows that as these models get more capable, the potential for harm scales exponentially.
We cannot just rely on post-training patches.
**Taylor** (1:31)
Right! Like if a regular kid can jailbreak GPT-5 to get a recipe for something toxic, the safety guardrails are clearly not as solid as they claim.
**Morgan** (1:42)
Exactly. And the fact that hundreds of people were actively seeking this information shows there's a real demand for malicious use cases. It is not just hypothetical.
**Taylor** (1:52)
Totally. Some of those prompts were looking for biological hazards. If the model is giving them actionable advice, that is a huge liability for OpenAI.
**Morgan** (2:03)
It raises huge questions about self-regulation. If OpenAI is downgrading risk ratings internally while these issues persist, we might need stricter government oversight.
**Taylor** (2:15)
Yeah, I think this is going to fuel the fire for AI regulation debates. But alright, let's pivot to something a bit more creative, though still mind-blowing.
**Morgan** (2:25)
Good idea, because my anxiety is spiking. What is this creative breakthrough you have got?
**Taylor** (2:31)
Oh, you are going to love this. It is from Black Forest Labs, the creators of Flux.
**Morgan** (2:38)
Wait, is this about Flux 3?
I saw some chatter about it on my feed this morning. What is the big deal with this version?
**Taylor** (2:46)
Dude, Flux 3 is a complete game changer. It is a multimodal flow model, meaning it learns from images, videos and audio inside a single architecture.
**Morgan** (2:59)
Okay, but is not everyone doing multimodal now? What makes Black Forest Labs' approach so different from, say, GPT 4 or Gemini?
**Taylor** (3:09)
Because it is the first Flux model to ship video, audio and robot action prediction from one single set of weights. It is all connected.
**Morgan** (3:20)
Wait, robot action prediction from the same weights as video and audio?
That is actually a huge deal. Usually, those are completely separate models.
**Taylor** (3:31)
Right?
BFL argues that no single modality gives the full picture. By combining everything into one flow model, the robot action stuff gets way smarter.
**Morgan** (3:43)
That makes sense theoretically. If a robot understands the sound of glass breaking and the video of it, it can act much faster and more accurately.
**Taylor** (3:52)
Exactly. It is like building a true digital brain that perceives the world more like we do.
The demos of it generating video and audio simultaneously are insane.
**Morgan** (4:05)
It sounds impressive, but I wonder about the compute requirements. Running a model that handles video, audio and robot control at the same time must be incredibly heavy.
**Taylor** (4:16)
True, but if they can optimize it, this could be the operating system for the next generation of humanoid robots. It is so cool.
**Morgan** (4:25)
If they can pull that off, it would be revolutionary. But I am still skeptical about how well it performs in real-time physical environments.
**Taylor** (4:35)
Fair enough, but the potential is huge. Speaking of models that are absolutely crushing benchmarks, we have to talk about Anthropic's new release.
**Morgan** (4:45)
Oh, the Claude Opus 5 news? Yeah, that one actually made me double-take. The numbers are almost hard to believe.
**Taylor** (4:53)
Dude, seriously, let's get into it, because the benchmark scores they just released are absolutely wild.
**Morgan** (5:00)
All right, lay it on me. What did Claude Opus 5 actually achieve on the ARC-AGI-3 benchmark?
**Taylor** (5:09)
So, it scored 30.2%.
4 more minutes of transcript below
Thousands of transcripts fetched by people building searchable podcast archives
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/YOUR_EPISODE_ID