Hermes Mixture of Agents DESTROYS Claude? artwork

Hermes Mixture of Agents DESTROYS Claude?

AI News Today | Julian Goldie Podcast

June 28, 2026

Hermes Mixture of Agents: Build a Model Council That Beats Single Frontier Models (Goldy Bench #2)Hermes released Hermes Mixture of Agents, which the speaker integrates as a “Hermes Council Engine” tab inside their Agent OS to run a panel of frontier models (e.g., Opus 4.8 and GPT 5.
Speakers: Julian Goldie
**Julian Goldie** (0:01)
So, Hermes just released Hermes Mixture of Agents, which is a powerful way to basically have a council of agents working together to get better outputs. And this is obviously based on the restrictions with Gpt 5.6, Fable 5 being taken down. It's like, okay, how can we achieve better levels of intelligence using systems where agents work together, and then you fuse the ideas to get the best possible outputs. And from what I've tested so far, this is a really powerful system. And I'm going to show you exactly what we've built with it, how it works, etc. How to use it, and also how it performs on Goldie Bench versus everything else. So let's test this out and see how it works. Now, if you want to do it in a case, how does this whole system work together? So basically, I call this the Hermes Council Engine. So you have a panel of frontier models merged by a chair. And I actually put it through my whole 42 task leaderboard. I've tested it with 42 tasks. You can check them out on Goldie Bench, as you can see right here for Hermes Mixture of Agents. And it actually ranked number two on the entire board, pretty much above every single model out there. The only one that's beaten it right now is Fusion. And also Opus 4.8 was outperformed by a long way. And I'll show you the comparisons in a second. So if you're wondering how does it perform on the benchmarks, it's doing pretty well. And I'll show you the examples in a minute. Now, if you look at the quote from News for Research on this, who created Hermes Agent, they said, Hermes Agent now exposes Mixture of Agent presets as virtual models, capabilities beyond the publicly available frontier. So if you want to indicate what is this? Well, basically, what we've done is we set up a council engine, which is a new tab built inside Hermes Agent, inside my agent operating system.
This is the setup right here. And so it's a new tab, we built inside the agent operating system. And basically, you can pick a few frontier models, as you can see right here. Any providers mixed together, I'm running Opus 4.8 and Gpt 5.5. And then when you ask a question, both models answer it privately at the same time. Then a third model, the chair, reads both answers, judges them, and writes one final answer that's better than either.
And we can see the actual outputs here in terms of the stuff created.
So the panel of models is the council, the chair is the aggregator. You have one question that you put in and one better answer out. So news research shipped this idea as a mixture of agents. I wired it into a tab, gave it a workspace and put my own leaderboard on it to see if it actually holds up in terms of performance. And basically the way this works is like a panel of experts would be one genius. So if you picture one brilliant person answering a hard question alone, well, if you now picture a panel, each expert writes their own take privately, a sharp chair reads all of them and gives you the best combined answer. So the panel would win pretty much almost every single time. So you have one prompt, Opus 4.8, for example, and Gpt 5.5 work together and then you get a chair that works and builds it and fuses it all together. Now, in terms of the performance, you might be wondering, okay, how does this hold up? So if we have a look at Hermes Mixture of Agents here, on the leaderboard is coming at number two, just below fusion. And if we pull up the answers here, you can just test yourself out yourself, right? So if you're thinking, this is not that good, or it doesn't look that great, et cetera, just have a look on this website and see what you think for yourself, right? You can make your own mind up or you can test it yourself. That's really why I created the Goldie Bench, is because we wanted to test all of these models on example prompts, rather than just listen to some theoretical benchmark in a lab somewhere about an AI model we can't even use yet. So that's why we created this system. And you can see here, for example, the stuff we've created is super nice, looks pretty cool, very visual. We've tested out on loads of different stuff. So for example, this is called the Dragon Realm, which is an open source project. As you can see, she's pretty cool, looks super nice.

11 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000774564297