No Moat: Closed AI gets its Open Source wakeup call — ft. Simon Willison artwork

No Moat: Closed AI gets its Open Source wakeup call — ft. Simon Willison

Latent Space: The AI Engineer Podcast

May 5, 2023

It’s now almost 6 months since Google declared Code Red, and the results — Jeff Dean’s recap of 2022 achievements and a mass exodus of the top research talent that contributed to it in January, Bard’s rushed launch in Feb, a slick video showing Google Workspace AI features and confusing doubly...
Speakers: Simon Willison, Alessio Fanelli, Travis Fischer
**Simon Willison** (0:00)
So yeah, this is a document which I first saw at 3 o'clock this morning, I think.
It claims to be leaked from Google. There's good reasons to believe it is leaked from Google. And to be honest, if it's not, it doesn't actually matter because the quality of the analysis I think stands alone. If this was just a document by some anonymous person, I'd still think it was interesting and worth discussing. And the title of the document is, We have no moat and neither does OpenAI. And the argument it makes is that while Google and OpenAI have been competing on training bigger and bigger language models, the open source community is already starting to outrun them, given only a couple of months of really, like really, really serious activity. You know, Facebook Llama was the thing that really kicked us off. There were open source language models like Bloom before that and GPTJ, and they were very impressive. Like nobody was really thinking that they were ChatGPT equivalent. Facebook Llama came out in March, I think March 15th, and was the first one that really sort of showed signs of being as capable maybe as ChatGPT. I think all of these models, the analysis of them tend to be a bit hyped. Like I don't think any of them are even quite up to GPT 3.5 standards yet, but they're within spitting distance in some respects. So anyway, Llama came out and then two weeks later, Stanford Alpaca came out, which was fine-tuned on top of Llama, and was a massive leap forward in terms of quality. Then a week after that, Vecuna came out, which is to this date, the best model I've been able to run on my own hardware. I've run it on my mobile phone now. It's astonishing how little resources you need to run these things. But anyway, the argument that this paper made, which I found very convincing, is it only took open source two months to get this far.
It's now every researcher in the world is kicking in on new things. But it feels like there are problems that Google has been trying to solve, that the open source models are already addressing.
Really, how do you compete with that? With your closed ecosystem, how are you going to beat these open models with all of this innovation going on? But then the most interesting argument in there is it talks about the size of models and says that maybe large isn't a competitive advantage. Maybe actually a smaller model with lots of different people fine tuning it and having these LoRa, stackable fine tuning innovations on top of it. Maybe those can move faster and actually having to retrain your giant model every few months from scratch is way less useful than having small models that you can fine tune in a couple of hours on laptop. So it's fascinating.
Basically, if you haven't read this thing, you should read every word of it. It's not very long. It's beautifully written.
If you try and find the quotable lines in it, almost every line of it is quotable. Yeah, that's the status of this thing.

**Alessio Fanelli** (2:49)
That's a wonderful summary, Simon. Yeah, there's so many angles we can take to this. I'll just observe one thing, which if you think about the open versus closed narrative, Emad Mostaque, who is the CEO of Stability, has always been that open will trail behind closed because the closed alternatives can always take learnings and lessons from open source. This is the first highly credible statement that is basically saying the exact opposite, that open source is moving, but then closed source and they are scared. They seem to be scared.
Which is interesting.

**Travis Fischer** (3:25)
A few things that I'll say. The only thing which can keep up with the pace of AI these days is open source. I think we're seeing that unfold in real time before our eyes.
I think the other interesting angle of this is to some degree, LLMs, they don't really have switching costs. They are going to become commoditized. At least, that's what a lot of people think. To what extent is it a rate in terms of pricing of these things? They all become roughly the same in terms of their underlying abilities. Open source is going to be actively pushing that forward. Then this is coming from, if it is to be believed, the Google or an insider type mentality around where is the actual competitive advantage? What should they be focusing on? How can they get back into the game?
When the external view of Google is that they're spinning their wheels and they have this code red and it's like they're playing catch-up already, could they use the open source community and work with them? Which is going to be really, really hard from a structural perspective given Google's place in the ecosystem, but a lot of jumping off points there.

39 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000611908357