**Julian Goldie** (0:00)
Run Hermes Agent Free Forever. So, Gemma 4 just dropped a brand new update, and it's now actually three times faster. Plus, there's a bunch of other improvements to it, including tool cooling, which is perfect for using Hermes Agent. So, we've actually already plugged it in to our Hermes Agent over here, and we've got a new profile for Gemma 4, and this is a free local model. The great thing about this is, the one where we're using Hermes Agent, it's private, number two, it's running locally, number three, it is free to run Gemma 4, and you can plug it in to Hermes Agent, and it can do agentic stuff. So let's have a look at this example. We actually said forward slash learn, and then we gave it a guide, and you can see here that it's actually broken down the guide and understood exactly how to use this skill in the future. So we're using the brand new feature, which is forward slash learn from Hermes Agent, and it worked, and it's very fast to reply, which means that it's better than ever to use Gemma 4 with this setup. So what has changed here? Well, basically Google fixed the agent skill.
So the improvements rolled out across the Gemma 4 family, and the tool calling patch is the one that turns Gemma 4 into a chatbot, you can run locally into an employee, aka an AI agent that you can run locally. That is the unlock, and it makes it way faster and consistent and more accurate on tool creation. So this is something I call the self-owned agent, which means you can run Hermes Agent free forever, and Google just fixed one thing that made local AI agents unreliable.
So the way that I actually set this up is we actually created a new agent profile. If you want to do that, you can just go to Manage inside your Hermes Agent Dashboard, and then you can scroll down here, click on Add Profiles, and basically you want to make sure that you add the model from Olama, which is Gemma 4 If you're not sure how to get Olama, you just go to olama.com, download it, and then you'll see Gemma 4, and you can run Olama from Gemma 4 there. Or the other option is you can actually run local models like Gemma 4 with MLX using Huggy Base and LM Studio. So why would we do this? Well, the thing with using agents is that they use a lot of tokens, but the problem is that that can get expensive. And so if you're using a local model, you can actually have Hermes agent with a frontier model like Cheapity 5.6 as the main brain of the operation, and then that can delegate sub-agent tasks to something like Gemma 4 Or if you're feeling adventurous, then you can actually have Gemma 4 as the main brain of your operations. You can have Gemma 4 as the main brain behind your AI agents, and then you can build and automate anything for free and privately using this system as well.
And so if you look at this system, basically everything just happens on your computer. You got Gemma 4 that can run through Olamo or LM Studio, and then you got Hermes, which is the hands. So Gemma 4 is the brain, Olamo is the engine room, Hermes is the hands. You might also say, you know, free local models not that great. I wouldn't say they're anywhere near, for example, Table 5 or anything like that, even if you've got, you know, DGX Spark or something like that. But if you want free and if you want local models that don't use up a lot of tokens, then you can use something like Gemma 4 The other way that you can get the most out of this is not just inside Hermes, but for example, we have a local setup inside our agent-operate system, where we can use Gemma 4 to build a code locally. We can preview what it built, and then everything is saved inside our workspace. So we have the preview tab where we can see what we've created. We've got the build section where we can see what we've built, and then we have the workspace to see what we've created.
You might also say, okay, is this actually good for creating stuff?
We actually tested it on a bunch of builds, including a website. The website was by far the most impressive design. I think it's because we actually gave it a custom skill to create the web design as well. It looks beautiful, right? It created this really nice website, as you can see, looks really clean, looks super nice. So Gemma 4 is getting better, and I've seen a lot of improvements rolling out from it recently, including MOX, which also makes it faster as well. And so the brain is literally one file, you know, then you've got Olamma, or you can run it fully locally as well. You can type a message to Hermes, and the brain thinks locally, it can use tools, so Gemma 4 can use tool use as well. And then it just works privately and locally for free as well. Also, the good thing is about Gemma 4, is that there's lots of different sizes. So depending on your setup, if you have a very lightweight setup, then you can use one of the smaller models. You'll see that ranges from a minimum of 6 gigabytes, which is super small, and that you can see that's on the MOX, or you can use the bigger models, for example, like up to 19 gig. And you can also check it out on Hugging Face as a model, and then get the setup from there too. The other really good thing that this is useful for is agent loops. So for example, if you've got big builds, for example, like the forward slash goal mode, using an agent loop with a local model, it can run all day, it could run on your computer, but it doesn't cost you anything to do that.
5 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000777138086