Topics: Technology, Business, Investing
**Josh** (0:00)
Meta has pivoted again for the third time, and now back to an open-source company. The story of Meta has been pretty insane. Two years ago, they were totally open-source, hell-bent on wreaking havoc on the entire marketplace of closed-source models. They were doing this through their model called Llama. And for those who weren't familiar, Llama was just kind of like this close to frontier open-source model that was expected to change the world in a meaningful way. It did not go as Zuck foresaw, and therefore they went and they closed off all of their model development and went closed-source just like all of the other labs. Now, they're going back. They're transitioning back to open-source, and as a result, they've spent a tremendous amount of money on this. So what happened in the last earnings report? After all the capex spend came in, Zuck just dropped a 14-page paper explaining exactly why they're making a re-entry into the open-source world. That entry is defined by a singular model that we're going to talk about right now called Muse Glimmer. And Muse Glimmer is a really impressive model that fits on your laptop, runs locally, and I think it's going to shift the world of open-source AI.
**Ejaaz** (1:01)
I think this is a great move from Zuck and Meta. And this is coming from probably the Meta's biggest bear in the past on this show.
**Josh** (1:10)
Yeah, definitely. We're professional haters. I should preface this with.
**Ejaaz** (1:13)
I'm a professional hater of Zuck and Meta for the last six months, but I love to see this pivot. And I'll explain why. Number one, Meta's entire philosophy around AI has always been open-source. And his simple reasoning behind that is he believes like everyone should get access to this thing, right? If everyone has access to this thing, there'll be no centralized core power for AI. There's no AI overlord going forwards. He then kind of pivoted on that, like you mentioned, and now he's back with not one, but two new open-source models. So what are the models? Number one, the headline is called Muse Glimmer. It's an open weight, open-source, 30 billion parameter model. Now, if you're not familiar with the sizes of these relative models, it's a pretty tiny model. It can actually fit and run on your MacBook. It's 20 gigabytes of data or RAM that it requires, which is significantly smaller than any other frontier model, which is like hundreds and hundreds and hundreds of gigabytes. And the way that they were able to achieve this is they took a bigger model and they made it smaller. It's something called quantization. Now, why did they create this model? Well, he believes or Meta believes that locally run private AI models should be the future of how AI models are dispersed amongst everyone. Right now, we pay subscriptions to get access to a model that lives in the cloud, that lives in a centralized company like Anthropic, like OpenAI. With this model, you can kind of run it locally, and that brings several different advantages that we've mentioned quite a few times on this show. The first one being it's cheap as hell. The second one being it can run on your private data, so it can become a lot more smarter and attuned to you specifically.
And the third thing is, it's very efficient when it serves you answers and results to your different prompts. So I think it's a good model, but it's certainly not frontier in any way.
**Josh** (2:59)
Yeah, it's far from frontier, and I imagine they're probably not really concerned about that. We mentioned this on previous episodes where there is this fight for the frontier, but now there are different frontiers. There is the frontier of intelligence, which Anthropic and OpenAI are kind of working on. Then there's the frontier of cost and efficiency per token. That's kind of where I see Meta placing themselves. In fact, they're fighting on multiple fronts at once, and it seems like Muse Glimmer is sitting at the very base of that because it is such a small model, it's so quantized, it's able to actually run on a local machine. I think the fully capable, largest version of this takes about 55 gigabytes of memory, which runs on my MacBook that I have sitting on my desk right now, and that's really impressive. This puts them in a race that not many other companies are competing on, which is just like that edge inference local compute thing that I think we can expect to see out of Apple Fair League soon. Like with the new Siri going to be running on local models, this is very much feels like a response to that, where now for the first time you're able to run this free inference that's fairly capable. It has a somewhat large context window. I believe a couple of hundred thousand tokens, I'm not sure the exact number, but a couple of hundred thousand tokens of context. It has pretty good outputs in terms of tokens per second, and it runs at a level that you would expect a Frontier model would have ran at maybe say 18 months ago. So it's not going to do any crazy, difficult, complicated tasks. But for something that can run on your machine, you could use anywhere in the world at any given time, and you can trust it to manage all of your data. So for example, if you have a machine runs a lot of sensitive information, you want to put health information through that, you might not want to give to a public facing model. This is a really easy way of doing that in a way that we haven't really had before. And I think that's kind of step one in this new strategy that Zuckerberg is going for in I guess you could call it like Scorched Earth 2.0, where why do you open source models? Well, you want everyone to go off, build on them to remove margins away from these frontier labs. And this seems like the first version of that, that people are kind of excited about.
24 more minutes of transcript below
Thousands of transcripts fetched by people building searchable podcast archives
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/YOUR_EPISODE_ID