**Julian Goldie** (0:00)
There is a brand new local model called North Mini Code that is super powerful for coding, and you can also get access to it for free. So this is available on Olamma now. So North Mini Code and Olamma are working together, and it's pretty powerful, as you can see right here. So if we look at the benchmarks, it's a little 30 billion parameter model that scores models four times its size according to the Artificial Analysis Coding Index. So you can see, for example, North Mini Code, which is performing at 33.4 here, is outperforming DevStrile, Mistrile 4, and Nemetron 3 Super, which is unbelievable. Now we've got it running inside our agent operating system already. So if we go to the local section here, we can build with it, and we have North Mini Code plugged in. You can also plug this into your AI agents, and that means it's free, it's local, it's private, it can work off-line, and it can actually build stuff. So for example, let's give it a little test run here, just as an example, build a colorful to-do list app.
And then we can control this with our voice, we can start building with it, and you can see that's working right here. And so this is local, it's ready to go, it can actually build stuff, and it's free to use, which is unbelievable. Now, if we actually go to the preview section here, we can see what we've previously built, so everything that we've built is saved, and then you can see that is quickly coding out here. Now, once that HTML is fully coded out, we can actually preview it, as you can see.
And if we test it, it actually works, unbelievable. So this is a free local model, just dropped, works properly, works offline, works with Hermes agent or whatever agent you want to plug into. And also we've got a workspace here where you can use your local models with this whole system. And it's really good and ready to go. So you might be wondering, okay, how does North Mini Code work? How does it perform, et cetera? Let's pull up the benchmarks right here. So if we compare it on benchmarks, it's not quite up there with Quen 3.6, but you need a good set up for Quen 3.6. However, it is outperforming Gemma 4, as you can see on Terminal Bench. It's also outperforming Gemma 4 on Terminal Bench Hard, SW Bench Verified is absolutely crushing Gemma 4 And it's right up there with Deathstral Small 2 And it's up there with Quen 3.6 as well, surprisingly. So it's an agentic coding model that's pretty powerful and easy to build with, as you've seen. You can even control your voice. You can build stuff with it. It actually works and it's ready to go. So pretty powerful stuff.
So let's talk about the local AI coding engine. This is a free coding model that lives inside your computer. It's fast, it's private, it's made for building, and five things make it work. So number one, it's a coder. So North Mini Code is trained for one job, which is writing software. Number two, it's a small body, so it's only 30 billion in size, but only a few experts actually fire up a word. So it's quick to run, as you saw. And it outscores models four times larger at coding, which is unbelievable. So you can outperform models four times its size because it's small, but it's a mixture of experts model. It has reasoning on, so you can switch on reasoning whenever you want to. It's free and it's under the Apache 2 license, works offline too, and you can build by voice as well.
So let's compare these side by side. You know, you could pay for a subscription, but that would get expensive, or you can use a tool like this to basically code unlimited. You've got a free coder built only for software. It beats models four times its size. It works at 92 words a second on a normal Mac with no limits. Nothing you type ever leaves your machine. You say the word and the whole app appears as you've seen today. And the result is a fast private coding engine that you actually own. That is unbelievable. Now, how does it work? Why is it so quick? So it only wakes a few experts per word. This is what makes it so fast. You can run this on a laptop too. So old models wake the whole brain for every single word, which means big models are quite slow locally. North is split into 128 experts and only eight of them fire for each word. So it's a giant brain with tiny effort. And that's how 30 billion parameter model runs as quick as a small one as you've seen.
6 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000773565861