**Julian Goldie** (0:00)
Today, we're going to be testing out a new local model called Laguna XS 2.1. And this is a model that seems to do pretty well. It's pretty fast and easy to use as well. We've already installed it and plugged it into our Agent OS. Show you how it's working in a second. And it just got updated today. So this literally just dropped. It's available on Hugging Face, and it's pretty powerful from what I've seen so far. Now it is quite a lightweight model. It's pretty chill to use. You can see how it performs right here on the benchmarks. So you can see, for example, Laguna XS 2.1. There was a previous version called XS 0.2.
And you can see how it performs versus Qwen 3.6. Now, if you're familiar with local models, you're probably familiar with Qwen 3.6, which is very, very popular, and North Mini Code, which I've already tested and done a showcase on that already. So if we have a look at them side by side for SWE Bench Verifier, bear in mind this is a coding model. It is a model designed for coding. You can see here that it's not far off Qwen 3.6.
So it's a little step up from XS 2, but it's pretty close to Qwen 3.6.
Then you can see, for example, North Mini Code, which was actually pretty decent when I tested it, it was quite fast as well, is being outperformed by Laguna XS 2.1. Now you might say at this point, Julian, I don't care about benchmarks. I don't pay attention to benchmarks. I want to see what you've tested. No problem. We're going to come on to that in a second. So don't worry, my friend. But I just want to show you how it looks side by side. You also might say, okay, well, how does it compare against like, you know, other models? So Claude Haiku 4.5, which is not that relevant now, but I could see why they've added it. It's obviously, you know, local models are not frontier models, right? So, you know, if you're hoping that local outperforms frontier, I would probably wait a long time for that. But if you want to see, okay, how does it perform versus something slightly more comparable, like Claude Haiku 4.5, you can see that it does outperform Haiku on SW Bench Pro. It also outperforms Gpt-Oss by a long, long way. So this is Gpt-Oss and this is Laguna XS 2.1.
So overall, it's an exciting model. I'm excited to test it and let me show you what we've built with it so far. So just to recap, it is from a company called Poolside and they've released Laguna XS 2.1. This is a new agentic coding model. So it's designed for agentic coding and terminal work. Now it has a 256k context window, which is enough for a local model to be honest. And if you want to know, okay, what did we build with it? Let me show you what we got from it side by side. So we created three different pages with it. This is the first one, which is kind of like a simple landing page. Honestly, I've done the same test with something like Gemma 4 And this actually turned out better than Gemma 4 So I would say for agentic coding, for stuff like building landing pages, for creating, for example, simple mini apps, this is not a bad alternative at all. I mean, it looks quite nice, quite clean on the page. Again, it's not going to be like Fable 5 level, but hopefully if you're into local models, you don't expect that either. Now we've also got another app here. This is like a to-do list app, so we can type in our to-do, and then we can add it inside here. Again, like if you look at this UI is not that nice. So you either want to lower the expectations on UI or train it on a really good skill on how to use AI because straight out of the box, it's not going to create something amazing. So if we have a look at this and we're like, okay, check out local models, add it. It actually works, which is great. We can delete it, we can add it, we can see what's added, what's done, click, complete, etc. So it actually works quite nicely.
Then we have this page as well.
So this is another page that we built out with the local model. Again, if you look at this, it's like, okay, this is something that I could create with Claude last year. But again, it's usually about a year behind when you're using local models. So is it fast? Can it code? Does everything that are built work? Actually work? Yes, absolutely. So it's not bad at all. And again, I'm running this. You might be saying, okay, Julian, what setup have you got here? So I'm running this on a Mac Studio M4 Max, 36 gigabyte memory, and it seems to run pretty smoothly, which is nice as well. Also, something to note here, not good at all for 3D models. Now you also might say, okay, well, how does it perform versus all the other local models? So this is brand brand new, just dropped less than 24 hours ago. We're going to plug it into our benchmarks on Goldie Bench, along with the newer version of Gemma 4 and MLX, it runs 90% faster. I've already done another tutorial on that. But basically, we'll be testing out the 42 different builds and seeing where it performs. You can see who's top of the leaderboards here. And if you want to check out what we're building with local models and how it works, et cetera, you can check out everything that's created inside Goldie Bench. And you can see side by side what we've created, how it performs, how good it looks, et cetera. And then just test out for yourself as well. Test this stuff out, don't believe the benchmarks.
6 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000775397041