Hermes AI Agent + Ollama: FREE + 1 Click Setup! artwork

Hermes AI Agent + Ollama: FREE + 1 Click Setup!

AI News Today | Julian Goldie Podcast

June 7, 2026

Run Google Gemma 4 Locally for Free in 1 Click (Ollama + Hermes Agent Setup)The episode shows how to run Google’s newly released Gemma 4 locally for free using Ollama and connect it to Hermes Agent in a single click/prompt, highlighting new quantization-aware training weights that reduce memory...
Speakers: Julian Goldie
**Julian Goldie** (0:00)
Ollama has just released a new free local model that you can run with Hermes in one single click. So you can see here, for example, Gemma 4 Quantization Aware Training weights are now available on Ollama. So basically what this does is it reduces memory requirements whilst maintaining the model quality. Why is that better? Because basically you can run Gemma 4, which just got released this week from Google. You can run it for free with Hermes Agent. And the cool thing about this is you can set it up in one single click. Let me show you how. So basically if you're using Gemma 4, you can see this was just updated a few hours ago. And if you want to run it with Hermes Agent, you can just copy this and then run it inside your terminal.
Now we've already plugged it into the agent operating system. So you can see Hermes here. And I just tested out, made sure it's working. Seems to be working fine, which is great. So you can see here it can use browser use, it's connected to Mac, so it can run things locally. And also it's got access to our memory.
Also, it can coordinate multi-agents and it has access to everything inside the terminal as well.
And so this model that we set up is Gemma 4 If you go over to the model section here, you can see we are running Ollama, launch Gemma 4 with our AI agent. Now, why is this useful? Number one, it's free. So if you're worried about APIs, you can set this up. Number two, this is quite a lightweight model. So for example, here, if we look at the size of the models, normally you might be looking at something that's 20 gig, for example, if you're setting up a local model. But they've actually got some really lightweight ones here. So for example, Gemma 4 12b. And then you also have these new quantized aware models, like for example, Gemma 4 Now, this was just announced a few hours ago, so you can see it from Google Gemma here.
What this basically means is that it reduces the amount of memory required whilst making sure it still performs well. Why is that important? Because if you're running an AI agent with local models, obviously it uses a lot of tool calls, it uses a lot of stuff. And if you have a slow setup, or if you have a model that's quite slow to respond, then it slows you down and the performance of the actual agent itself. So you can change this. You could also, for example, have a different model, especially if you want to be on a free model, you could have a different model that's running for free. So for example, you can actually get Gemma 4 for free on Open Router, or you could select Nematron 3 Super, which is available for free on Open Router. Additionally, you could use NVIDIA Nematron. Also a little tip here, if you when you're setting the main model inside your dashboard with Hermes Agent, if you go to change and then type free, you can set free models like. Now, if we type Ollama, we can switch between models here as you can see.
So we have Gemma 4 ready to go right there. What you could also do is you could use this for a sub agent. So you can actually change the agent you use for other tasks, like smaller tasks to a local one because it doesn't require much brain power. And then for the main model, you could select that as something a bit more powerful. And if you wanted to keep it all free, again, you can use like News Portal and just change that over to a free model, like Step 3.7 Flash or Nemetron 3 Ultra.
And so you have an AI brain that lives on your machine, set up with Hermes Agent and now it's private, it's local. If you don't have Wi-Fi, you can still use AI and you can plug it into AI Agent. And also it's easier than ever because you can install it in one click. You can also plug this into Cloud Code, CodeZap, OPCodex, OpenCode. Would I recommend that for Cloud Code? Probably not. But I would recommend it for OpenCode. I think that could work pretty well. So it's depending on what you're running it on as well. So, for example, I'm running this on a Mac Studio, so it seems to handle Gemma 4 fairly well. And so you might be saying, OK, what is Gemma 4? This is an open source project from Google. It's actually designed for agentic tasks. So they basically announced earlier this month, Gemma 4 is our most capable open model yet, built to run locally, built to run privately, with no telemetry and free for any use, personal or commercial, which is awesome. And so it runs for free. You can download it. It runs on your machine. And so the old ways like paying for subscriptions, being worried about data, not being able to use it offline, etc. The new ways you can download Gemma 4 once, takes about 20 minutes to download, depending on how fast your internet is, then it's free. It's open source. It runs on your machine. It can use Hermes or OpenCore. It's agentic as well. It works offline. It's yours. So you keep the data as well.

6 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000771572740