How to Run Hermes FREE Forever! artwork

How to Run Hermes FREE Forever!

AI News Today | Julian Goldie Podcast

July 3, 2026

Run Hermes Free Forever: Gemma 4 MLX Update Makes It Up to 90% Faster (Ollama + Apple Silicon)Julian demonstrates how to run Hermes for free using a new Gemma 4 update with Ollama on Apple Silicon via MLX, claiming up to 90–95% faster local performance compared to before.
Speakers: Julian Goldie
**Julian Goldie** (0:00)
Today, I want to show you how to run Hermes for free, forever, with a new update to Gemma 4 That actually makes it 90% faster. So this is a new update that just dropped for Gemma 4 with Ollama on Apple Silicon using MLX. And so with MLX, you can run models 95% faster using Gemma 4 Gemma 4 is a local free model from Google, and you can now plug it into Hermes in like one single click and run free models forever. So let me show you an example. This already plugged it into the HNOS over here. And then if we go to the different profiles that we have, I usually just have one for each different API that we use. And the one that we've got set up running for free is Gemma 4, as you can see right here. Now we just make sure it works, it responds pretty quickly. We can also build stuff with it locally. So you can see, for example, here, we said like build out a to-do list app, and it build out in a couple of minutes, and you can see how many tokens it's running on as well.
And then we can see the preview over here as well. Everything that we built with this is directly here. Now the beauty of this is that it's way faster than it used to be. It used to be so slow, it's basically unusable. Whereas with this system, you can now run it. Now you also might say, okay, Julian, this is just for Mac. Julian, I don't have a silicon, whatever.
What you can actually do here is you can use a free API on OpenRouter. An OpenRouter have the 31B model available as a free API that you can plug into Hermes as well. So if you think this is only for people on Mac, or you can't use this if you don't have Apple Silicon setup, you actually can. You just use the free API instead. So don't worry about that. But either way, you can run Hermes agent free. The difference, of course, is that if you run a local model instead, then it's all private, it can work offline. You don't need the internet, you can use it on a flight, whatever. And so the great thing about this is agent workflows can now run on a free local model. Hermes is one of the most powerful agents in the world, and you can now run it with these free local models. Let me show you another example. So if we pull up Hermes over here, and we say, okay, forward slash learn, and let's just plug in the details of a recent tutorial that we've created like this.
So we can take this example, we can go to type forward slash learn, and then we can get Hermes to start thinking and working on that task. Basically, when you use forward slash learn with Hermes agent, it can read the tutorial and it can add that as a skill so it never forgets it again. That's just running the background with a free local model, which is beautiful. So you can have autonomous agents, they're just getting stuff done in the background, which is unbelievable when you think about it.
So with Gemma 4 fast enough locally, I used to avoid Hermes agent with agentic stuff but now with this system, we don't need to avoid it so much because we can just have Hermes running in the background. This might not respond to you within seconds. Bear in mind, it's working on quite a complex task there and it's set up a skill set. It would normally even with a frontier API that would take a few minutes. But the point here is you can have it running in the background. You don't need it to respond to you straight away. It's not a chat bar, it's an agent. It can build stuff for you 24-7 whilst you're off doing something else.
It's great. We can set a loop.
We can walk away, it can build stuff. It can work on the Kanbab. We could have multiple agent profiles building with Gemma 4 using Hermes and we can run it free forever, which is absolutely amazing. You can see an example of how to set this up. If you're wondering, how do you run this? How do you use it, etc.
Well, it's pretty simple. The way that this works is you can just make sure that you use Ollama with Hermes and you can run the terminal command for it. So if you go to Ollama, make sure you have the latest version. Once you've done that, go to Models, then you're going to go to Gemma 4 and the model that you want to use is important to note here is these are the normal models, the old ones and these are the new ones. So you can see this was just updated and this is MOX you want to be using. So you can run the MOX models and those are the ones that will be 90% faster. And so the great thing about this is you can just run in a loop without you 24-7 using this free local model.

7 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000775393974