**Sarah** (0:05)
Hey, listeners, welcome back to No Priors. This episode marks a special milestone. Today is our 100th show. Thank you so much for tuning in each week with me and Elad. And it's been an exciting last couple of weeks in AI, so we have lots to talk about. Why don't we start with the news of the hour, or really the last month at this point, and DeepSeek. Elad, what's your overall reaction?
**Elad** (0:28)
DeepSeek is one of those things which is both really important in some ways, and then also kind of what you'd expect would happen from a trend line perspective. And I think there was a lot of interest around DeepSeek for sort of three reasons. Number one, it was a state-of-the-art Chinese model that seemed to have really caught up with a number of things on the reasoning side and in other areas relative to some of the Western models, and it was open source. Number two, there was a claim that it was done very cheaply. So I think the paper talked about like a five and a half million dollar run is sort of the end. And then lastly, I think there's this broader narrative of who's really behind it and what's going on and some perception of mystery, which may or may not be real. And as you kind of walk through each one of those things, I think on the first one, you know, state-of-the-art open source model with some recent capabilities built in, they actually did some really nice work. You read through the paper, there's some novel techniques in RL that they worked on and that, you know, I know some other labs are starting to adapt. I think some other labs had also come up with some similar things over time, but I think it was clear they had done some real work there. On the cost side, everybody that I've at least talked to who's savvy to it basically views every sort of final run for a model of this type to roughly be in that kind of dollar range, you know, $5 to $10 million or something like that. And really the question is how much work went in behind that before they distilled down this smaller model. And my sense is everybody thinks that they were spending hundreds of millions of dollars on compute leading up to this.
And so from that perspective, it wasn't really novel. And I think that sort of 20% drop in NVIDIA stock and everything else that happened as news of this model spread was a bit unwarranted. And then the last one was just sort of speculation of what's going on. Is it really a hedge fund? Is something else happening? Like, you know, thoughts a little bit, well, speculative. There's all sorts of reasons that it is exactly what they say it is. And then there's some circumstances in which you could interpret things more broadly. So that's kind of my read on it. I mean, what do you think?
**Sarah** (2:25)
Yeah, I think it's interesting sort of the delayed reaction to it. But to your point, it's also like what you might expect, especially given historical precedent with like GPT-335 and then ChatGPT. So like DeepSeek v3, like the base model, big AI model pre-trained on a lot of Internet data predict next tokens. Like that was out in December, right? And NVIDIA stock did not crash based on that news.
So I think it's just interesting to recognize that like people obviously do not just want raw likelihood of next word in a streaming way and the work of post-training and making it more useful for human feedback or more specific data, like high quality examples of prompts and responses just like we've seen with the chat models, like ChatGPT, the instruction fine tuning that made this such a breakthrough experience like that really mattered. And then the, as you said, the like narrative violation release of R1 reasoning model, as a parallel model to like OpenAI's O1, I think that was also the breakthrough moment in terms of people's understanding of this.
**Elad** (3:32)
Well, it's also like 20 years of China-America technology dominance narrative, right?
**Sarah** (3:37)
Yes.
**Elad** (3:37)
Like I think, I think it was also kind of this zeitgeist around US versus China. You know, worst, you know, the West is far ahead. And so, you know, will they ever catch up, et cetera? And this kind of showed that Chinese models can get there really fast. But I do think the cost thing was a huge component of it. And again, I think cost may have been in some sense misstated or misunderstood, at least.
**Sarah** (3:55)
19 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000689937579