**Julian Goldie** (0:00)
There's a brand new update from Open and Fusion that basically allows you to achieve, according to this, according to this, Fusion allows you to achieve fable level intelligence, but at half the price. So you can see the charts right here in terms of how it performs on benchmarks. And this is pretty interesting as a model. So you can see here, they're basically Fusion on 100 hard research tasks. And there are panels of models that consistently outperform individual models. So what this is, is basically like you can create a panel of three different models that work together to get better outputs and also to use less tokens. So you can achieve beyond frontier performance with frontier panels. And this is very different to using one individual model. It's a panel of agents that work together. And you can also have this is really interesting. This way it gets fascinating. So you can have panels of budget models, right? Cheaper models that can surpass frontier models. And obviously that's cheaper, right? It uses less tokens, it uses less powerful APIs. So you can see an example of the tests right here. By testing different combinations of models, they found roughly three quarters of the lift the Fusion provides comes from synthesis and one quarter from diversity. So you can see an example of how they perform right here. So what this means essentially, this is very interesting. So you've got Opus 4.8 solo. So Opus 4.8 working alone. Then you have the benchmark score with Opus 4.8 and Opus 4.8 working together. Now if you have Opus 4.8 and 5.5, you get even better results from the working together. But you can use cheaper models. So you can use even free APIs will be quite interesting to test for this.
So what's the most interesting out of all this is that the budget panel was actually comparable with Claude Fable 5 in performance.
So if you have like a panel of Gemini 3.5 Flash, Kimi K 2.6 and Deepsea V4 Pro fused together as a panel of agents working together, that would be Solo 5.5 and Solo Opus 4.8.
So bear in mind, these are not as powerful, they're cheaper models, but they would be frontier models because they're working together. And actually landed within 1% of Fable 5 on the intelligence tests. So you might be wanting to locate how does it work?
How do you use this? So when you send a prompt to Fusion and you can do this via API, it fans out to a panel of models in parallel, each with web search and bash tools enabled. So what happens after that is a judge model reads every response and extracts consensus points, contradictions, partial coverage, unique insights, anything they might have missed. And then you can actually use this, for example, inside a chat, like you can see here, and you can choose which models you have working together. Now you can switch between this. You can select, for example, a budget panel or you could have a quality panel or you could have a custom panel.
And this is really interesting because now you can have multiple agents working together inside a panel. So let's test this out. If you've got the quality section here, we can plug in a prompt like so. And then from here, it's going to start using all three models at the same time with a judge model coming in later. Really, really interesting stuff. So these are now generating and they're just working together separately. And you can have up to eight models in parallel working together. So you could have like Opus, Gemini.
Obviously you can't use Fable anymore. That's gone. But the difference here is like you could achieve potentially Fable level intelligence with one judge that fuses and gives you one answer. And that's the cool thing as well as like, you don't have to check five different answers at the same time. The judge fuses the models and then gives you one answer back. You can have eight models in parallel working. So we've got that working over here, as you can see, and see score higher on benchmarks. And this was interesting as well. There's a full breakdown on it here in terms of what they found and how it works. And these are the tests that we've done. Well, not we've done, they've done. So as a fusion, Fable 5 and GPT 5.5 synthesized by Opus 4.8 as a judge got the highest score on these benchmarks. But if you look at these benchmarks, Opus 4.8, GPT 5.5 and Gemini 3.1 Pro synthesized by Opus 4.8 scored within one of Fable 5 with GPT 5.5. Now, if you look at Claude Fable 5 solo, that scored 65.3%.
8 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000772894420