Gemma 4 is now 90% FASTER + FREE + Local 🤯 artwork

Gemma 4 is now 90% FASTER + FREE + Local 🤯

AI News Today | Julian Goldie Podcast

July 3, 2026

Gemma 4 Just Got Up to 90% Faster on Mac (Ollama + MLX)Google’s Gemma 4 has a new update that makes it up to nearly 90% faster on a Mac when run through Ollama using MLX, enabling much faster local model performance.
Speakers: Julian Goldie
**Julian Goldie** (0:00)
Today, we have a brand new update from Gemma 4, from Google's Gemma 4, and it's now nearly 90% faster on a Mac, with Ollama, using MLX. And this means that essentially, you can run Gemma 4 way faster than ever before. I've been testing out, I'll show you what we've built out in a second. We've also plugged it into our agent operating system, so we can build and automate anything that we want. We can preview over here. Honestly, it is way, way, way faster, which is fantastic.
We're liking that. And we're going to show you exactly what we've built with it today, how it works step by step, and how you can use it too. So the great thing about this is that, honestly, local models up to this point have been pretty bad and pretty slow. And I'm on a Mac Studio. So if I'm on a Mac Studio and it's not that great, then that is a bit of a nightmare, my friend. So what we have now is we have MLX with GitHub. So MLX is a way to run local models directly on your Mac, right? It's designed for a Mac because usually, like if you're running local models without MLX, they're super slow. And then what we also have is we have Ollama and on Ollama, you can use Gemma 4 So I just want to be 100 percent clear before people get confused or whatever. Just to be very clear here, Gemma 4, free model. You can get it on Ollama, MLX, free open source project. You can get it on GitHub. And so with this combination, that means you can build and automate whatever you want using these three models and they're way faster and way better than they've ever been before, which is fantastic. So you can see here they've actually added MLX models directly to Ollama. So this means that you don't even have to get MLX directly. You can just download Ollama, which is free, and then guess what? You get Gemma 4, which is free, and that's basically it. And then you just make sure you install one of the MLX models. Now, once you've done that, then you can plug it into a system like this. So we have the agent operating system, and this basically takes, you know, it's plugged into my memory. It's got all of my agents plugged in as well. So we could run, for example, Gemma 4 directly for free with something like Hermes agent. We actually did that right here. We've got Gemma Speed, we called it, because it's 90% faster, and we've got that plugged into Hermes, which is perfect. And so this is a pretty powerful system where, you know, local models are actually good, actually useful. So Gemma 4 just got faster, up to 90% faster. I will show you what the actual tests on my own, tests show you when it comes to speed. I think from what I tested, it can be up to 90% faster. That doesn't mean it's 90% faster for everything. So, for example, when I was testing out, I think for most of the stuff that we tested out, it was about 60% faster. I actually got Claude to measure it for us, 60% faster. But still that's pretty good for a free local model. You know, we can have it switched on and it's ready to go. And so this is a big, big jump really. And I'll show you what we've built with it in a second, how it works, et cetera.
And by the way, this just dropped. It was only on the 1st of July that this update dropped. You can see the comparison here. So we, it was running at 50 tokens per second previously. And now with MLX, it's running at 95 tokens per second. So big jump and improvement right there. And you can see the on Hermes when we've got Gemma speed running or Gemma 4, whatever you want to call it.
We said working and it's working right there. So it's good to go. Pretty useful. We can use it with Hermes. We can plug it into FreeClock code. We can plug it into our local model builder. And then what we've done here is we've set up a system inside our agent OS where we can ask it to build something.
And then from there, we can preview it. And we can also see it inside our workspace, which is super useful too. And we actually separate up our models. So for example, when we were testing out the local model Quable, we have the preview over here. When we had, for example, Gemma 4 being tested, we have that over here, right? Which is great. So if we go back into the build section here and we say create a to-do list app, and then we just hit enter, that will start thinking locally and building locally with Gemma 4, using Gemma 4, MLX, and then it will actually show us a preview once that's completed. So we can have that working in the background. We could have it running 24-7 if we wanted to as well. And so the good thing about this as well is like, everyday builds, the little stuff, the stuff where you need a lot of work going on, etc.

7 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000775400768