**Julian Goldie** (0:00)
So we have a brand new update from a Japanese AI lab called Sakana, and they've created something called Sakana Fugu. I probably totally mispronounced that, but this is designed to be a full multi-agent orchestration system, accessible via a single model API. And this has another model inside it called Fugu Ultra that matches the performance of Fable and Mythos, apparently delivering frontier capability, as you can see right here. Now, basically what this does is it, it's quite similar to Fusion, if you've seen Fusion, what it should have is basically a mix of models working together, and then they create an answer, right? So instead of just having like one model like Fable 5, you would ask a question, and then this has a panel, which is a multi-agent panel API that competes head on, right? And then it synthesizes the answer, and you get one answer with, I mean, it literally just dropped an hour ago.
So go easy on me guys. But what I will say is we've already tested out on three different things, and I would always say test out yourself. I will explain more on that in a second. So we can see an example of a website we built here. Actually looks super nice. I'll show you how this compares on Goldie Bench with everything else that we've created recently. But the website itself looks super nice. So that was one example. We also have, for example, this Maze game, as you can see right here.
That turned out super nice. And again, I'll show you how this compares with other models, because we always use the same tests and the same prompts with every single model. But it's pretty powerful stuff and actually looks really good. And then we've got this one as well, which is pretty mind blowing. This is like a, you know, a simulation of the galaxy. And the quality of the outputs here is super nice, super nice stuff. So, you know, if you're looking for a Fable 5 level model, or you're looking for Fable 5 level outputs, this could be it. But again, I would test out yourself. I've already tested out myself, as you can see here. And I get the idea of this because we've actually used something else similar recently, and it's called Fusion. Right. Now, if you've never used Fusion before, it's a similar sort of idea. You have multiple models as a panel, and then they go for a judge which fuses it, and then you get one answer out. Now, if you're wondering how does Sakana perform on benchmarks, let's have a look at this. Let's pull this up right here. So we've got Terminal Bench, and you can see Fable 5 scored 80.4 versus 80.2 and 82.1 versus Fugu. And on all the benchmarks here, it's pretty much outperforming Fable 5 SW Bench Pro, Fable 5, destroys both Fugu and Fugu Ultra. But on most of the benchmarks, they're pretty even or it's being outperforming. So you can see, for example, Live Code Bench here, Fugu Ultra 93.2, 92.9 and 89.8. So pretty impressive model, pretty powerful stuff, really interesting idea. We've already tested it, like I've said before. Now, if you're wondering, okay, how does this compare against everything else? Let's have a look head on. So let's, for example, compare this versus Glm 5.2, and we can see the same tasks side by side. So this is Glm 5.2, this is Fugu Ultra. Let's have a look here. So you've already seen this demo, and then let's have a look at the version from Glm 5.2, which is still a little bit buggy, as you can see right here when you compare it on benchmarks, and pretty difficult to navigate and move around.
Now, if we have a look at the next example here.
So this is the website. So Glm 5.2 versus Fugu. Fugu looks super nice.
As you see, nice animations, nice colors, nice UI, etc. And then if we compare that versus Glm 5.2's output, which still looks nice, it's just not quite got that same touch. It just doesn't look quite as nice. So side by side on the outputs here, I would say that Fugu is winning on the benchmarks. Again, test this stuff yourself, see what you think. Let's have a look at the Galaxy example here.
So this is the living galaxy, the living spiral galaxy where we can zoom in, we can zoom out, we can move this around, etc. And then if we compare that versus Glm 5.2, looks totally different, right? Totally different. I would say which one is more interesting, which one is more beautiful.
I would go with this version right here. It's just way more interesting to use on a deeper level. And the outputs are pretty amazing, pretty inspiring stuff. I've also got more tests running in the background. So we're testing this out. Let's compare it versus Opus 4.8 as well. So this is Opus 4.8 output and this is Fugu's. Like which one looks a lot more interesting, a lot more powerful? For sure it's this one, right? More interesting, better design. I will say just even Fable 5 was not that great, a UI. But if you compare them side by side, this one looks a lot nicer.
8 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000773830367