**Julian Goldie** (0:00)
So today, I'm going to be showing you Qwable 27B Coder, which is a new coding local model. And I'm going to show you exactly how it works. Now, we actually built some interesting stuff with this, far more than what we usually create with local models. So that impressed me in the first place. Here's, for example, a landing page that we created locally using Qwable. And it was pretty easy and simple to create. It moves in the background. It looks better than like 99% of those sort of local websites that you usually create.
I wouldn't say it's anywhere near like Frontier level, but at the same time, like you can create some pretty cool stuff better than what some of the things I've built with other models. And I'll show you some example comparisons in a second as well. We'll look at how it compares versus other local models I've tested recently, like for example, Onif or Quifos.
And I do think these local models again, better and better. And this is one of the best ones I've seen recently. So this is a cool little game we built out. Here's another one.
The thing that I will say here is that if you look at this, it kind of feels like last year's frontier, if that makes sense. So it's not like going to start competing with Opus 4.8 or Fable 5 tomorrow, but it can build some pretty cool stuff. And this is actually better than I expected it to be when it comes to to build it out with these local models. So how does this work? Well, essentially, what we built here is Qwable 5 27B Coder. It's open source, it's free to run, and it just dropped on Hugging Face. Pretty good so far. Just got updated end of June, and it's got a Qwen 3.6 27B base. So that is the base for creating this. So it's free to use, free to use locally. And we actually ran it with Apple MLX, which is a free open source setup just for using local models. So we can run LLMs from Hugging Face using the system. And that's how we ran it.
You can't get it through Ollama.
And if you want to do how it performed on the benchmarks here, yeah, it did pretty good, right? I would say in terms of local models, as far as they go, it's right at the top of the leaderboard from everything we've tested out recently. So for example, recently we've tested out Gemma 4 12B, Qwifos 9B or Onif 1.0. And I would genuinely say like the quality of stuff that we got from Onif is nowhere near the same level as the quality of stuff we got from Qwable.
Loving these names by the way. So it's actually my favorite local model so far. I'm on a Mac Studio, Apple M4 Max, 36 gigabytes of memory. When I actually run and create stuff with this model, I can definitely feel like the whole setup runs a bit slower, but at the same time, it can run in the background whilst I do other things. So it's not like going to completely slow down your whole setup. So we can still, for example, like run Cloud Desktop whilst this was running before. Now also something that we did is we plugged it into our agent operating system so that we can generate live previews whilst we're building with this stuff. So for example, if we give it a command, well, our local engine inside the agent operating system runs with Qwable so we can run this on free models now like Qwable 5 And then when we say build something out, it will actually preview it so we can see what we've created and then open up the preview. And everything that we create is plugged into our workspace. So for example, that 3D Dragon game that we just talked about, that is available to preview inside our workspace right here. And then we can come back to everything that we've created, which is pretty cool. Also inside the workspace, we can open it up inside a new tab or we can get the code from it directly. But that's basically a really cool way to build with local models, preview what you've created and then run it on free models as well. Now, obviously Qwable kind of a reference to Fable 5, but that kind of oversells it. So it's based on Qwen 3.6 27B, which is a strong base model from Alibaba. And it's got a fine tune as well. So it's basically Qwen 3.6 27B with a flashy Qwable Coder jacket. And then if you're wondering how to run it, if you can't run it from Ollama, you would run it with Apple MLX if you're running it on a Mac. But yeah, the stuff that I built was pretty nice. The only problem was that it's a lot slower. So if you were comparing it to, for example, like Qwifos 9B, Qwifos 9B is way, way faster than Qwable 27B. So that's something to be aware of as well. It's like, yes, it builds better stuff, but it's going to slow you down. However, it's top of our leaderboards when it comes to Goldiebench and the local models we're testing out. And this is something I'm just going to keep building out over time. So you can see the comparisons and see how they perform. The cool thing as well, like you can run your Agent OS now on free local private models. You don't need Wi-Fi to use these as well. And they're just ready to go whenever you need them. Now if we have a look, for example, at the same task with Gemma 4, Gemma 4 12B Coder totally failed on that task. So if we click on this, for example, this is the same game and it just, it just didn't work. It didn't work at all.
3 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000774767332