**John Coogan** (0:00)
I need some soundboard. Here we go. Yes, today on TBPN. We're talking about Model Mayhem. Everyone's launching new models. Slow summer, but not for the AI race. You got XAI unveiling Grok 4.5, the first model built specifically for coding in AI agents, developing collaboration with Cursor. Talked about it a little bit yesterday, but we have some more benchmarks, some more discussion on the timeline about where this model fits in on the Pareto frontier.
Also, why it might be outperforming so well on Cursor bench.
Lots of debates there. Meta announced Muse Spark, a new agentic coding model with Mark Zuckerberg. Returning to X for the first time in basically a decade.
Three years ago, he posted one joke post about launching threads, but he has not been an active user, but the AI vortex sucked him in and he's got a post.
**Jordi Hays** (0:52)
Oh, I think he's an active user, John.
**John Coogan** (0:54)
You think so?
**Jordi Hays** (0:55)
He's just not an active post. He's just not an active contributor.
**John Coogan** (0:58)
You're calling him a lurker.
**Jordi Hays** (0:59)
I'm calling him a lurker.
**John Coogan** (1:00)
You're calling him a lurker?
**Jordi Hays** (1:01)
I'm calling him a lurker. I think he's absolutely glued.
**John Coogan** (1:04)
You think so?
**Jordi Hays** (1:05)
I think so.
**John Coogan** (1:05)
You really think so?
**Jordi Hays** (1:06)
I think so.
**John Coogan** (1:07)
I feel like, I don't know, so busy, so much other stuff going on. I feel like most people are on that level.
**Jordi Hays** (1:14)
The busiest people I know are not active on X, but they are on X a lot.
**John Coogan** (1:22)
Sometimes, but there's a different class of person.
**Jordi Hays** (1:24)
You can just quiz them.
**John Coogan** (1:26)
Screenshots come to them via Slack or via text message, because they have a team that's monitoring the timeline and then it's delivered. This is the important stuff.
**Jordi Hays** (1:35)
They're calling him Mark Zuckerberg.
**John Coogan** (1:39)
But the other big news, OpenAI just released GPT 5.6. Let's go. Let's give it up for a new general purpose model with expanded coding and agent capabilities alongside GPT Live, which we talked about yesterday, a new real-time interactive voice experience. Reactions are great to 5.6. Bunch of interesting details here. You had, people have been identifying that while there is a frontier and there are just a few companies that are actually on the frontier, the frontier is spiky and they have different flavors to them and different reasons to pull different tools off the shelf.
People are drawing analogies between Fable 5 being some recluse genius and 5.6 being a collaborative co-worker that you love chatting with or something like that.
**Jordi Hays** (2:33)
I said, I don't know how else to describe it, but Fable 5 is like Kendrick on Good Kid, Mad City and 5.6 Soul is like Chief Keef on Finally.
**John Coogan** (2:42)
Now it makes sense to me.
**Jordi Hays** (2:43)
Thank you for bringing it up. So I just wanted to put it into 2010 hip hop terminology.
**John Coogan** (2:49)
Really really clear there. Thanks for clearing that up.
**Jordi Hays** (2:53)
The funny thing is that will be very explicit for like 100 people in the whole world.
**John Coogan** (2:58)
The most interesting benchmark to me has always been Arc AGI v3. We've interviewed the team over there many times and had a lot of fun understanding what goes into that benchmark. And 5.6 sole scored a massive 7.78% which is tiny.
Considering that the whole point of Arc AGI is that a human should be able to get 100% on it and basically any human. So it is a true test of AGI in this sense of, you know, can you give this test to just actually anyone, not, you know, the crazy math projects, the crazy hard programming projects, hacking, all of that stuff is very economically valuable, of course, but there's a more interesting question where, you know, when there's less of a spiky frontier and there's just this question of what is something that anybody can do that AI can't? Because we've been searching for those and the Arc AGI team has done a fantastic job building out these puzzles that AI has historically struggled with. Arc AGI, one, the model sort of climbed, two became a little bit more complicated, and now three, we're starting to see glimpses of progress, although 7.76% isn't 99%, we're nowhere near saturation, but it's still a huge jump. Opus 4.8 had 1.5%, so GPT 5.6 sole is showing more generalization, more spatial reasoning, more puzzle-solving abilities. So fun, fun stuff. The blog post is also very, very fun because it includes games. I'm a big fan of the GPT 5.6 launch games. I got immediately sucked into the sailing minigame, which is very high fidelity but also delightful to actually play. Should we play it? Yes, we should definitely play it. Yeah, Salt Wind.
19 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000776186080