Topics: Technology
**Nathaniel Whittemore** (0:00)
Today on the AI Daily Brief, a massive AI leadership shake up at Google. And before that, on the headlines, Meta drops two new models and a coding harness. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI.
All right, friends, quick announcements before we dive in. First of all, thank you to today's sponsors, KPMG, Rackspace, Blitzy, and HyperAgent. To get an ad-free version of the show, go to patreon.com/aidailybrief, or you can subscribe to Apple Podcasts. And if you want to learn more about sponsoring the show, send us a note at sponsors at aideailybrief.ai.
Meta continues its comeback KQuest with the release of MuSpark 1.2 and Mu's code. Alongside the Twin Model release, they are releasing their first coding harness as well. Meta described MuSpark 1.2 as a coding-focused update to the 1.1 version which was released in July. This is the first model that Meta has trained in a harness, improving its agenda capabilities in that environment. The results look like a pretty strong coding model on the benchmarks. It scored 82.9% on Terminal Bench 2.1, placing it between Opus 5 and GPT-56 Terra. On DeepSui, it scored 59.3%, placing it behind Opus 5 and GPT-56 Terra, trailing by around 5 points. Meta chose not to compare MuSpark to the frontier models, likely because it's not in the same size class as Fable 5 or GPT-56 Sol, and the model appears to be designed to be cheap and efficient as a daily driver, rather than taking on the larger models on the benchmarks.
Artificial Analysis had similar findings. The model scored 54 on the AA Intelligence Index, placing it behind Opus 5, GPT-56 Tera and Kimi K3, tying it with Grok 4.5 and putting it a few points ahead of GLM 5.2. AA also wrote that Spark 1.2 is quote, among the most cost-efficient models at its intelligence level. It cost 40 cents per task on their benchmark run, which gave it a similar cost to intelligence ratio as Grok 4.5 and GPT 5.6 Sol turned down to medium effort settings. Its run was around half the cost of Kimi K3, further reinforcing the idea that every Chinese model is not just some incredibly low-cost wonder. Mu Spark 1.2's run on the AA index was around half the cost of Kimi K3 with results in the same ballpark. AA also noted that the three-point overall improvement was almost entirely down to agentic performance. The update delivered a big jump on GDPVAL, making it the sixth-highest ranked model behind Opus 5, Fable 5, Quen 3.8 Max, GPT 5.6 Sol, and Kimi K3.
On the harnessed side, the biggest thing besides Meta actually bringing a coding harnessed to market is subagents. In his launch thread, once again on Twitter, where Mark Zuckerberg has been spending a lot more time recently, Zuckerberg wrote, Muse code runs specialized background agents that stay active your whole session, so they build up context over time instead of starting from scratch on every task. When a job is big enough, it fans out to separate subagents working in parallel in isolated work trees. Your working copy is never touched. In testing, we had it build six features for a game simultaneously with no collisions. Meta saw very strong performance for long-horizon tasks with this architecture. During testing, they deployed the model to a kernel optimization task, and the model successfully ran for 24 hours, executing more than 1,000 tool calls, and delivering steady improvements throughout the session. Meta is also selling Muse code as suitable for professional work due to its auditability. The harness logs every tool call and code edit, and has the ability to use these logs to restart midway through a task if it crashes. Now, aside from the release, Zuckerberg also teased parts of the future roadmap. He wrote that larger and more capable models are on the way, as well as hinting that Muse code might be open-sourced.
Now, so far, if you dig around, you can find both positive and negative responses. I would say overall, the steady drumbeat of each sequential release getting a little bit better from Meta and them getting closer and closer to relevant again continues with this set of releases. Teasing what will be the subject of our main episode, Hader wrote, How quickly the tables have turned. Meta is starting to look like Google, moving fast, shipping AI products, and finally building momentum. Meanwhile, mighty Google suddenly looks like last year's Meta. Confused, reactive, and somehow watching everyone else move faster.
Now, this was not the only Meta story in the news. Not to be left out of the current trend, Meta's agents have also escaped containment and hacked into third-party systems. Meta said that during cybersecurity testing for MuSpark 1.1, the model left its sandbox and exploited a vulnerability to break into systems owned by another unnamed company. A Meta spokesperson said the issue was a misconfigured sandbox provided by security evaluation partner Irregular. Irregular, by the way, was also involved in the incidents reported by OpenAI and Anthropic, with those companies stating that the sandbox failed to quarantine the model away from the open internet. Meta said that the hack was carried out, quote, in a manner similar to previously reported instances with other companies. They're still conducting an investigation and said they will provide a full report once they have all the facts. But a spokesperson for Irregular confirmed that the Meta incident involved the exact same sandboxing issue previously disclosed by the other labs.
29 more minutes of transcript below
Thousands of transcripts fetched by people building searchable podcast archives
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/YOUR_EPISODE_ID