Loop Engineering Makes Hermes 10X Better! artwork

Loop Engineering Makes Hermes 10X Better!

AI News Today | Julian Goldie Podcast

June 20, 2026

Hermes Loop Engine: Automate Quality Control With a Builder + Judge SystemThe script introduces a Hermes loop engineering system designed to improve output quality from Hermes agents, especially when using free, cheap, or local models.
Speakers: Julian Goldie
**Julian Goldie** (0:00)
So today, I want to show you a new Hermes loop engineering system that we've put together, that has basically helped us improve the quality of using Hermes, and also to get more, especially out of like free or cheap models or local models when we're running Hermes agents. And the way this works is basically you set the bar, so you tell it what done means, the builder acts, so it creates and writes a draft, the judge grades it, you know, and you can use a free model to actually grade the outputs and see if they're good. And if there's any issues, so if it's not very good, then it loops back around and it keeps looping back ground with basically a Hermes agent, the quality controls and checks everything until it's finally done. Now, what does this mean? This means you set the bar once, you walk away and you get way better outputs from Hermes and you have to babysit less, you have to monitor less, and also you automate the quality control of your outputs with these agents. So a builder writes it, a separate free judge can grade it out of 100 It fixes itself round after round until it passes and you stop being the loop. Now, if you want to see an example of that in action with Hermes agents, we've actually built out the loop engine inside here. So you can see some examples of what we built down here. And basically, you just define what your definition has done and then you type in your starting point here. This is inside my agent operating system in the AI Profit Boarding, and then we have the judge, the grader. You can also change the models that you want to use to grade it. So, for example, obviously, if you had like, I mean, this would be wild. But if you had, for example, Fusion as the builder and then the judge as Fusion as well, which is like five models working together. Wow. The outputs would be absolutely outrageous. It would take a bit more time, but that's just an idea and an example. So you could use, for example, like three builders, but then you can have a frontier model that actually judges it and gives feedback. Then you can also set how many rounds the loop goes round four. So as you can see here, you set the by the builder axis, and it keeps going round. Now it could go round forever, but you can just say, okay, just go around for like four loops. Then once it's done, it's done. So it's a really powerful system that we can use to get stuff done. You can see some examples of what we actually built in terms of apps and games and that sort of thing just for fun.
This just basically loops around with a free model creating this stuff, and a free model on the judge as well, that was grading this stuff and giving it a score. And so pretty amazing when you think about it, it works as a loop to just get better outputs. And I think this works so well is because like, you know, AI models, they hallucinate, sometimes they don't understand UI, so they can't get very good outputs in terms of like the website or the app that they design and build. And so this is just a way of automating the quality control. I'll give you an example, like you could ask AI for like a cold email, then you read it, it's not good enough. So then you have to tell it was wrong. And then you try again, you read it again and round after round, it's just you in the middle every single time, the judge, the note taker, the one giving feedback and pressing go again. And so like sometimes by like five rounds, you have something that's okay, but you've wasted a lot of time and energy managing your agents. Whereas with this system that we've set up, you automate it with Hermes and you don't need to worry about it again. And so now you can write down what done actually looks like, a builder writes it, a separate judge grades it hard, and lists what's wrong. It fixes it round after round whilst I'm doing something else, and you just come back to a finished result. And that's basically what we've built out with this system. And so, you know, if you are watching this and you're thinking, do you know what my AI agents, the stuff that we get back from them is not that good, we need to call it to control it better. Definitely look at this Hermes engineering loop system because it's super powerful. And just promise yourself one thing right now, you know, you'll finish this guide and actually set up one loop before you sleep tonight, just one, because the moment you stop being in the middle of every AI cycle, the way you work changes for good. Now the people sitting still are just still reading drafts all day and going back and giving feedback to the agents is super messy. The people who make this transition today get their time back, be one of those people commit to the transition commit to taking action. Because this changes how you work with AI forever when you think about it, because you're basically engineering a positive feedback loop where everything improves and the quality of everything improves as well.

9 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000773486206