**Nathaniel Whittemore** (0:00)
Today on the AI Daily Brief, how we're grappling with AI advancements when many of us can't even judge the new capabilities coming online. Before that, in the headlines, a new model that seems to have an impressive cost profile. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI.
All right, friends, quick announcements before we dive in. First of all, thank you to today's sponsors, KPMG, Blitzi, Robots and Pencils, and Airtable. To get an ad-free version of the show, go to patreon.com/aidailybrief, or you can subscribe on Apple Podcasts.
If you are interested in learning about sponsoring the show, send us a note at sponsors at aidailybrief.ai.
We have a bunch of interesting stories today. We've got a new model that's capturing a bunch of attention, more models hacking out of containment. But first, we come back to the story of the world's most famous AI hedge fund, which is apparently down but not out, as portfolio manager Leopold Aschenbrenner briefed clients on the situation. Now, shortly after I recorded Friday's episode discussing the blowup of Situational Awareness' public market portfolio, a letter to investors explaining the status of the fund was leaked. In the letter, Aschenbrenner explained that the portfolio had suffered a severe drawdown throughout July, exacerbated by quote-unquote adverse trading against stocks known to be held by the fund. Then on Wednesday night, Aschenbrenner wrote that the fund decided to take decisive action, selling off a portion of their public portfolio to remove all leverage. This move, he wrote, allowed the fund to protect their private market positions, which are generally believed to be heavily concentrated on anthropic. Dispelling some of the rumors, Aschenbrenner wrote that the fund, quote, was not shut down, liquidated, or transformed into a private-only fund. Most importantly, he added, we took the steps that were necessary to fight another day.
Aschenbrenner closed the letter with the claim that their unaudited numbers had the fund down 67% for the month, but still holding on to a net year-to-date performance of plus 80%.
Now, this report triggered a gigantic argument, largely split between the AI and finance factions on X. TBPN led their Friday show with the news and proclaimed, rumors of his demise are greatly exaggerated. Some trumpeted that the fund was still up 80% for the year after a nasty drawdown. Greybeard Investor and Constant noted that the unlevered semiconductor index is up 60% for the year, and the levered version is still up 3X despite the drawdown, questioning just how good 80% really is in this market.
And while there was skepticism about whether the fund could recover, certainly some were throwing their hats in the ring already. Uber successful AI angel investor reposted Leopold's note, declaring that he asked to invest in the fund for the first time. Even on the finance side of X, many noted that there's a long history of notable investors having an early blow up and continuing with a storied career. Even Citadel CEO Ken Griffin, who bought the distressed portfolio from Situational Awareness last week, suffered a 55% drawdown in 2008 and clawed his way back to become a titan of the industry.
Now, there is still a ton of speculation about what the Situational Awareness portfolio actually looks like post blow up, but it seems pointless to speculate when we can just wait for the next round of SEC reporting. For now, it is clear that Leopold's story is not over and he will continue to be a player in this market.
Next up, the latest in our stories of small models making a bid to undercut the next generation of ultra-large models, Deepseek has announced their new V4 Flash model. On the Artificial Analysis Intelligence Index, the model scored 50 That is a 10-point jump over the previous iteration of V4 Flash and 6 points higher than the larger Pro version. Against the field, Flash is firmly in the middle ground, tied with Gemini 3.6 Flash, and just 1 point shy of GLM 5.2 and GPT 5.6 Luna. There is a big gap, of course, between V4 Flash and the Frontier models, but this is not a model that's designed to compete on the Frontier. Instead, this could instantly become the most cost-efficient model available if performance lives up to the benchmarks. V4 Flash logged just 3 cents per task on the AI benchmark run, which is an incredible efficiency against comparable models like GLM 5.2 at 59 cents per task and Metamu Spark at 36 cents per task. It even beat GPT-56 Luna, which came in at 5 cents per task, for only a slight improvement on the benchmarks. This version of V4 Flash also managed to use 12% fewer tokens compared to the previous iteration, and logged a pretty significant jump on GDPVAL-AA, suggesting a significant improvement on agentic use. Now while people's initial impression was to be incredibly impressed with the price drop, their first results were perhaps a little underwhelming. Martin Casado, who had just lauded the model in a previous post, tweeted, hmm, Deepseek V4 Flash results aren't great for me. K3, on the other hand, is quite impressive. I wonder if we're actually hitting model size limitations on quality. Others had better experiences. Bookworm Engineer wrote, Initial thoughts about Deepseek V4 Flash. It feels like sorcery. I have been testing Deepseek Flash on all my work that I did with Fable and Kimi K3. My short verdict, I cannot believe this model is real at this size.
25 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/YOUR_EPISODE_ID