AI Companies Still Haven’t Delivered on Their Biggest Promises artwork

AI Companies Still Haven’t Delivered on Their Biggest Promises

The AI Daily Brief: Artificial Intelligence News and Analysis

August 17, 2026

Anthropic CEO Dario Amodei says the strongest criticism of AI companies is that they still haven’t delivered the enormous benefits they’ve promised—and that no amount of marketing can substitute for real results.
Speakers: Nathaniel Whittemore

Topics: Technology

**Nathaniel Whittemore** (0:00)
In a recent podcast appearance, a prominent investor said that he had heard from multiple sources inside Anthropic that Dario Amodei and other leaders in that company felt that at some point in the future, they might be the only company left. It would just be them, governments and the rest of us.
Now, these comments on that podcast kicked off quite a firestorm of discourse about Anthropic and their role in AI and what their beliefs actually meant for the industry. It also generated that rarest of phenomenon, an appearance on social media from Anthropic CEO, Dario Amodei himself. In his response post, Dario discusses his real views on regulatory capture, what he thinks the real root of AI's trust problems with people are, and what he thinks could actually address those trust problems in the long run. So, did people find it enlightening, convincing? Did anyone's opinions actually change? And what does the whole conversation say about the AI discourse and Anthropic's place in it?
The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI.
Welcome back to the AI Daily Brief Headlines Edition, all the daily AI news you need in around five minutes. We got ourselves quite a Monday here, so strap in. First up comes a new model and one that is sure to kick off a lot of debate, ZAI has dropped GLM 5.3.
Now, you might remember that when ZAI released GLM 5.2 in June, it helped kick off this new wave of concern that we're still in right now, that Chinese open-weight models were closing the gap. Now, part of that was timing. The release came during the period where Fable 5 was locked behind government doors and OpenAI was delaying 5.6 for the same reason, but the feeling of Chinese models nipping on Western Heels compounded with the release of Kimi K3 the following month. Still, as has happened every other time, even acknowledging what these models are really good for, there has been a sense that they still are ultimately behind the frontier in pretty meaningful ways. So where did that leave us with Glm 5.3? Well, 5.3 is built on the same base model as 5.2, meaning it's not some massive multi-trillion parameter model. Still, ZAI claims that they've made some big advances purely by scaling reinforcement learning. On coding, Glm 5.3 scores 28.3% on Terminal Bench 3.0. That puts it around 5 points behind the frontier with Fable 5 and GPT 5.6 sole, but 11 points ahead of Kimi K3. The results on DeepSui were less impressive, scoring 66.9%, which puts the model half a point behind Kimi K3, three points behind Fable and six points behind 5.6 sole. For agentic use, 5.3 scored state-of-the-art results on Automation Bench and GDPVal, just slightly inching out the US frontier. Overall, it looks like that for high-level use cases like running agents and coding, 5.3 remains a bit behind the absolute state-of-the-art, but has squeezed a lot of performance out of a mid-sized model and may, in some cases, have overtaken Kimi K3.
Now, ZAI said improved cyber performance was one of the key focuses for their reinforcement learning run. They noted that during the hugging face attack, cyber defenders had been forced to use Glm 5.2 because the guardrails on frontier models had rendered them useless.
In a WeChat post, ZAI said, If the powerful attack ability is spreading, the defensive ability cannot be limited to a few closed-source model companies. That makes their performance gains on cybersecurity benchmark CyberGym a bit more interesting, jumping seven points from their predecessor to actually overtake Fable 5
Now, just to be super clear, because of course we're already seeing some people freak out about a Mythos level cyber model released as open source, that is not what these benchmarks are saying even if you take the benchmarks at face value. There is a difference in being Mythos or Fable level at finding vulnerabilities, and being Mythos or Fable level at then doing something about it and or autonomously executing cyber attacks. ZAI is also taking a phased approach to the release, testing the model with trusted partners before publishing the full weights. Now at the time of recording, we don't yet have the full benchmark run from artificial analysis so we don't have either a full barrage of independent tests, nor do we have a great gauge for token-adjusted cost. On a per-token basis, Glm 5.3 is less than a tenth of the cost of Fable or 5.6 Sol, and one-fifth the cost of Kimi K3. In limited testing, some users are finding that Glm 5.3 costs around two-thirds of Kimi K3 or Grok 4.6 for the same task.

27 more minutes of transcript below

Thousands of transcripts fetched by people building searchable podcast archives

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/YOUR_EPISODE_ID