Why Frontier AI Labs Are Terrified of Open-Source | Ahmad Osman
MTS
September 12, 2026
Ahmad Osman joins MTS to discuss running AI models on hardware you control, the evolution of local inference capabilities, and why open source model development matters for individual agency and data sovereignty. Turn ideas into software people love.
Speakers MTS, Ahmad Osman
TopicsNews
MTS (0:01)
Hello, everyone, and welcome back to MTS. Today, I am joined by an exciting guest, Ahmad Osman, who is the founder and CEO of Osmantic, which builds infrastructure for companies that want to run AI on hardware they control. And I'm excited about this episode because local AI used to mean accepting a major capability downgrade and exchange for control. But that is starting to change. Downloadable models have gotten dramatically better, and hardware is becoming increasingly more capable of running them. So today, Ahmad actually ended up bringing to the studio NVIDIA's new DGX station. I think we have a little pan here. Look at this.
Ahmad Osman (0:41)
Yeah, we got the DGX station in house. This is an awesome piece of machinery, by the way. It has a total of 750 something gigabytes of coherent memory. 252 out of that is an HPM3E, which goes at 7.1 terabytes per second. I use it for running GLM 5.3 Flash and Quinn 3.8 Next Flash right now. And it's giving me a lot of thousands of tokens per second at very good quality, like intelligence quality at home.
We can't connect this to the bar, unfortunately, because it requires like a 20 amps, which requires a little bit of a specific setup.
MTS (1:21)
And it's really hot today, so.
Ahmad Osman (1:22)
It's really hot in here today, so we're bashing on that, but this is an amazing beast of machinery.
MTS (1:27)
Yeah. No, Ahmad, I think one of the things that made you so popular on X is not only were you talking about local models, but a lot of your posts include actual pictures of the hardware, how you're actually running them, and you actually get to see what these look like. So I've actually never seen a GX station in person, like I've never been this close to it. I've only seen it through my phone screen and through your posts. And right before we went live, we were talking about you might actually be one of the most famous local AI influencers out there.
Ahmad Osman (1:59)
Thank you for saying that, but that is not the goal. I started in 2023 just trying to write about inference at scale and trying to teach people about how to use these things at home, like LLMs. And it became clear to me that within the right timeline if everything falls in place right, we would be able to run like very smart and efficient models at home and we'd be able to like, you know, get to the frontier intelligence at hardware that we own without giving away our own data. Like, you know, we've just seen the whole thing with OpenAI and like, you know, the mathematics like theory that's being proven and whether they're training on codex stations or not. And we, there's a lot of like, you know, talks about intellectual proprietary like, you know, data and stuff like that. And to me, that the only way to bypass that or like, you know, not worry about that is to actually and itself renting intelligence on it.
MTS (2:48)
Yeah.
Ahmad Osman (2:49)
And that comes with hardware and compute.
MTS (2:51)
That comes with hardware and compute. And so I want to go into a little bit of the history behind local AI models because local AI, even earlier this year, I would say, was like this hobbyist thing. It's like, okay, that's really cool for the mountain man who wants full control over their models. But for the regular day consumer, this just wasn't something they cared about, or it just wasn't something that was physically possible with the type of hardware, and it often meant having way worse capabilities. And so where are we today? And how did that start out? Yeah.
Ahmad Osman (3:21)
So if we look a few years ago, we had Lama 2, and then we had the capabilities of something like GBT-3 that looked quite far away from each other. But what we've been seeing is that, in back-per-parameter, and I've had talks about this, and I'm actually speaking about it more at the Open Source AI Summit, so there's more to come on this. But the basic idea is that every four months, we're catching up, we're finding that smaller and more efficient models are catching up to Frontier Intelligence, and the hardware that you own at home is being able to process talk-ins more intelligently and more efficiently. Like, you know, the whole idea, and I think I've mentioned this before to you, but like, GBT 4 used to require a data center. Right now, on an iPhone, you can run a Quinn 3.5, 4 billion-barometer model, and it would run, it would actually provide you with smarter outputs, it would provide you with more capable intelligence on a phone. So that should tell you something about the hardware, like, you know, the hardware that used to cost something like $500 a few years ago, like an RTX 3090 right now is being sold for $1,500. There's a reason for that, and it's because, like, that's more than MSRB. And it's because, basically, models are being optimized and architectures are being built in a more efficient way that that hardware that is so old is more capable now. So the idea becomes what, you know, what's that hardware going to be able to do in a year from now? That's the bit that I have been making on local AI for quite some time now. Wow.
33 more minutes of transcript below
Thousands of transcripts fetched by people building searchable podcast archives
Fetch the whole transcript
The demo key returns a sample episode in full, no card needed:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090Markdown with the speakers named, for your notes, your knowledge base, or anything that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000789180893