Sakana: BETTER Than Fable 5? artwork

Sakana: BETTER Than Fable 5?

AI News Today | Julian Goldie Podcast

June 25, 2026

Sakana Fugu Ultra vs Fusion vs Claude Opus 4.8 (42 Builds + Goldie Bench): Is It Fable 5-Level?
Speakers: Julian Goldie
**Julian Goldie** (0:00)
So, Sakana Fugu Ultra has been out for a few days, and this is, according to benchmarks, able to achieve Fable 5 level intelligence. Now, does it actually match that standard? How does it perform? Where is it up to? Is it better than Fusion as well? I'm going to cover all of that in this video today. I'm going to show you what we've built with it. We've actually built out 42 different things with Fugu Ultra. So you get a really comprehensive idea of how powerful it is, whether it's worth using, the best ways to use it. And we've tested it on Goldie Bench as well. I'll show you the benchmarks and the leaderboards in a second in terms of how it performs versus everything else. Also how it performs side by side and what we've built with it so far. So we have tested it absolutely relentlessly using our system here. So we have Sakana Fugu plugged into our agent operating system, like you can see. And then what we can do with this is we can build cool new stuff of it and then test it outside by side. The only problem was that when it first got released, which was just a couple of days ago, you were kind of limited by the number of tokens you could use. And so the problem with that is like number one, it was super slow. And number two, you couldn't really test it as much as you can without getting like locked out for five hours. So we've built 42 different things of it today. I'm going to just guide you through what it is, how to use it and also how it performs versus Fable 5, Fusion and Opus 4.8. So let's get straight into this.
We ran 42 builds to find out and we used the same 42 prompts with three models, Fugu Ultra, which is from Sakana.AI, we've got Fusion, which is from OpenRouter. Both of these on benchmarks are claiming to achieve Fable 5 level intelligence. And then we have Opus 4.8, which is the only or sorry, the most powerful model you can get from Claude right now. And we'll compare these side by side with our Goldie benchmarks. And we've tested it, you know, with real builds that you'll see in a second right here. By the way, if you're wondering, okay, what is Sakana Fugu? How does it work? So it's basically a way of orchestrating multi agents together, right? So you have Sakana Fugu, you have closed and open models, and that's an LLM pool that basically fuses and answers together. So the idea here is like if you have multiple agents working together, multiple models, closed and open together, then you tend to get better results than just one single output from one single model. And so on their benchmarks, they claim to achieve better level intelligence than Fable 5 So you can see the benchmarks right here. And this is compared against Mifos Preview, for example. And you can see Fable 5 here too on the benchmarks. And on a lot of the benchmarks from this Japanese AI, Fugu, which is a tiny little AI lab, they are outperforming Fable 5, right? Now, let's see how it performs in reality on the actual builds that we've created right here. So on the build, I'm going to be 100% honest with you, on our leaderboard with Goldie Bench, you can see here that Fusion is actually top right now. It's actually crushing everyone else. Fugu Ultra is not doing too badly, but just have a look at the builds yourself today and see what you think, and see what you think about in terms of what you can get out of it. Now, one thing to note as well is like, if you're using Fusion, which is from OpenRouter, follows the same idea, use multiple AIs together.
With Fusion, the difference is that you pay per API. With Fugu, you use a subscription. And so with the subscription is a flat plan, with Fusion, you pay per token. So that's one of the biggest differences. Also, Fugu Ultra is from Japan, whereas Fusion is from OpenRouter, which I think is the US.
And also, with Fugu, if you're on a flat plan, you can obviously run out of tokens, just like you would with a Claude subscription. Whereas the difference with Fusion is that with Fusion, you don't run out of tokens because you're paying per API, if that makes sense.
So let's test them out and see what we've got here in terms of builds and how they perform. So first of all, we've got the solar system test right here. And you can see that we have Fugu Ultra, Fusion and Opus 4.8. So we tested all three together. So this is the version from Fugu. This is Fusion and this is Opus 4.8. Now, if I had to pick one, I'd probably go with this one. I think it's the most sophisticated and elegant version. I do quite like the output from Opus 4.8. I would say that's coming in second. And then Fusion comes in last on that particular example. Let's have a look at the next one.

9 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000774238995