Fusion DESTROYS Claude? artwork

Fusion DESTROYS Claude?

AI News Today | Julian Goldie Podcast

June 23, 2026

Fusion by OpenRouter: One-Prompt Game Builds, Panel-of-Models AI, and Benchmark Results vs Opus 4.
Speakers: Julian Goldie
**Julian Goldie** (0:00)
So we have a new model out from OpenRouter, it's an API called Fusion, that is designed to achieve Fable 5 level intelligence. And today I'm gonna show you what I've built with it. Here's one example, and it is blowing me away in terms of how cool this is. Bear in mind, this was created with one single prompt. All of the stuff that I'm about to show you is created with one single prompt. Let me show you another example here. So this is pretty cool, this is fun to play, it's kind of like a 3D game where you go through a maze. And then let's have a look at this one. This is a Dragon Realm light you can see, pretty cool, fun to play, etc.
For open world game, this is kind of like a basic, you know, break the blocks game, but it seems quite fun to play as well. And I'll compare it against all the other stuff. So we've run the same prompt on many other models, and I'll show you the difference in a second. And this one is really cool as well. This is kind of like a neon racer game, as you can see.
Super futuristic, really interesting to play and a lot of fun. So let's run through what Fusion is first of all and how it works. And then we'll come on to how it performs versus other tests and where it scores on the Goldie Bench benchmarks. Let's get into it. So this is Fusion. If you're wondering how Fusion works, essentially what you have is five models that work as a panel, and then their answers to the same prompt go through a judge who fuses the answers, critiques them, is kind of adversarial with them, and then gives you one single answer. So the way this works is really, it's one chart because you use the API once, that comes up with five answers, that the judge then turns into one answer, and then everything that you saw a second ago was created with one single prompt, which is pretty crazy in itself. Now, in terms of the benchmarks, how does it actually perform? By the way, just want to say here, I've got Goldie Bench and I've shown you the benchmarks and I'll go deeper into those in a second. These are benchmarks from OpenRouter. Test it yourself. See what you think. Don't just listen to benchmarks. And you can see the announcement from OpenRouter here. So obviously, Fusion came out just after Fable 5 got taken down. And on the tests, basically when you're running multiple models together, the idea is multiple minds are better than one mind. So when they work together like this, you get better answers. And you can see how they perform. So it's designed to achieve Fable 5 level intelligence or higher. Now, if we look at the tests from OpenRouter, you can see how they perform. So Solo Fable 5 does pretty good, but GPT 5.5 with Opus 4.5 synthesized, aka you've got the judge model with Opus 4.8, looking at the answers and fusing them together, that tends to work the best. Now obviously, Fable 5 and GPT 5.5 synthesized by Opus 4.8, that is goated. But obviously, you can't get access to Fable 5 right now, so it's sad times, but you get the point. Now, the other thing as well, that if you are running an API, it can achieve Fable 5 level intelligence on benchmarks whilst being 50% of the cost.
So when you think about that, it's pretty wild.
And basically one API fuses the best output of multiple models. That's how this works. Now, in terms of the benchmarks, we've tested 42 tasks with Fusion so far. Draco Deep Research scores 69 with a budget panel. So bear in mind, you can have a budget panel, you could have a premium panel. And what I mean by this, you could have cheap models running side by side, or you could have expensive models running side by side. And obviously, if you use the expensive models, you're probably going to get better outputs.
But either way, it should be Fable 5 Now, you wouldn't use this for like everything, right? If you're using Fusion and a panel of APIs and a panel of models working together, you're probably just going to use this on the big tasks.
This wouldn't be like a day-to-day API that you run, but it's built some pretty awesome stuff. Here's another example.
So this is like a pool game emulator that actually looks really good. And we created like a voxel dash game, as you can see right here. And again, like the outputs are super nice, it's just, it's fun. A lot of the stuff that it's creating is like fun, looks really nice. We've used it on websites and stuff like that as well, but you know, obviously, I like to show the visual stuff that's fun to play with. Now, how does it perform versus everything else? So if we have a look here, for example, we've got Fusion versus Opus 4.8, and we can see side by side how they performed in terms of the actual outputs. So this one is Fusion, and this one is Opus 4.8. So let's have a look at Opus 4.8 first of all.

5 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000773818590