Maple AI: FREE Local Model on Your iPhone artwork

Maple AI: FREE Local Model on Your iPhone

AI News Today | Julian Goldie Podcast

August 5, 2026

Maple AI: FREE Local Model on Your iPhone
Speakers: Julian Goldie

Topics: Marketing, Business

**Julian Goldie** (0:00)
Today, we have a brand new update from Maple Preview. Maple Preview is a new model open source 20B A1B ternary weight reasoning LLM state of the art needs weight class. And it's actually very, very quick to use. So for example, if we actually go inside our local agent operating system, like you can see, and we say like, let's say, for example, build out a beautiful snake game.
It's actually very, very quick to respond. Look how fast it runs. Now, bear in mind, like I've run, for example, Gemma 4 in the past using local models. I've tested all of the latest local models that are actually worth testing on a Mac Studio. And most of them are terrible or run slow or slow my whole setup down. This is really, really quick. We can preview everything we've built. We've got everything saved inside our workspace. It's running for free and local. We can switch off the wifi, continue using it. And it's right up there with Quen 36B. So let me show you an example of this running side by side. This is the announcement, by the way. And the difference here is that it's really focusing on speed. So there's a speed quality frontier for local models. And that's really the game that's being played. So, for example, you want quality, but you also want it to run fast. Usually the better it is, the slower it gets. So, for example, if you look at Gemma 4, it's right in the middle there.
This is the average performance, which means the quality is up, and this is the speed. So Gemma 4, E4B, E2B, both in the middle. If you look at, for example, LFM 2.5, pretty fast model actually when you use it, that runs fine. Now, if we look at Maple Preview, according to their benchmarks. Now, I've tested it myself. It is pretty decent.
It's not the best in the world at coding. I'll tell you that for free. But it is pretty fast when you run a local model. Its average performance is way up there, and its speed is up there. Now, this is what it's all about, really. So if you look at models in the same caliber, like, for example, QWEN 3.6 27B, which is probably like the one that I hear about the most, but it's way too slow to run on my setup, for example, like a Mac Studio. QWEN 3.6 27B is a lot slower. The other option is 1-bit Bonsai 27B as well. That's considered a decent model, which is basically QWEN 3.6 27B. That could take like 5 minutes to respond to a simple question, depending on where you run it. So if you look, for example, a Maple Preview, they're looking at running it on an iPhone. That's the other legendary bit about this is like, it's focused on smaller devices, it's focused on mobile. And I think this is a future that's coming where we can run local models for free on a mobile device, which right now, if you look at the comparison in the speed, it's not really possible with something like Binary Bonsai 27B, AKA QWEN 3.6 27B.
So it's designed to be very light, but at the same time run pretty nice. Now, here's another example. So this is running on a MacBook Pro where it can be a lot more autonomous. So they've said it autonomously decides to remember details. And that's the other big advantage here is that it can adapt, it can improve, it can learn. Now, they've actually compared it against Claude Sonnet 5 with the same scenario. I mean, when I've tested it on coding tests, it's nowhere near the same level as Sonnet 5 So I'm not even going to entertain that idea, my friend, but you can see where it goes on the performance benchmarks. Now, you can test it out on their website. If you just want a quick test, it's available at chat.deepgrey, or you can get the open weights directly on Hugging Face. Now, I've already plugged it into my Gentic Operating System, which you can see over here, and we have the local section. The other great thing about that is we can switch in and switch out any model that we want to, and basically code for free. Also, I quite like the way that it responds.
So it feels quite nice when you're using it.
It gives you some detailed responses. It certainly gives better responses than a lot of the other models that I've tested. Let's say, for example, build out a beautiful landing page for an SEO agency, and then we can leave that running over here. It will run locally for free. And also the great thing about this is is that it codes really fast, but we can also use our other agents inside the OS whilst that's running in the background. So for example, you could have LFM, which is another model I was testing earlier today, and that works really nicely with Hermes Agent. And then you can power your whole agentic operating system using free models. You've got the local builder over here and you have LFM over here. Now you might also wonder, okay, what does ternary weight reasoning mean when it comes to local models? This is something I actually learned today. So normal AR models store every connection as a precise number with many decimal places. Maple stores each one is just minus zero or plus. So just three symbols. That's what ternary weights means. So writing the brain with three symbols makes the file tiny and the mouse fast, which means a model that would normally need around like 38 gigabytes can fit into five gigabytes and your Mac's chips, for example, can run through it way faster.

3 more minutes of transcript below

Thousands of transcripts fetched by people building searchable podcast archives

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/YOUR_EPISODE_ID