Emergency Pod: Reinforcement Learning Works! Reflecting on Chinese Reasoning Models DeepSeek-R1 and Kimi k1.5 artwork

Emergency Pod: Reinforcement Learning Works! Reflecting on Chinese Reasoning Models DeepSeek-R1 and Kimi k1.5

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

January 25, 2025

This episode explores the groundbreaking advancements in AGI from recent releases of two Chinese reasoning models: DeepSeek's R1 and Moonshot AI's Kimmy.
Speakers: Nathan Labenz, Erik Torenberg
**Nathan Labenz** (0:05)
Welcome to The Cognitive Revolution. Today, I'm going to do a walkthrough of everything that I am learning and understanding and taking away from the latest Chinese reasoning model releases that have come out this week. Perhaps not coincidentally, both R1 from DeepSeek and the new Kimmy reasoning model from a company called Moonshot AI were released on Trump's inauguration day. And we now have two Chinese models. DeepSeek is out. The weights are open source. You can download them. But the Kimmy paper from Moonshot came a little later in the day. Honestly, shades of OpenAI and Google kind of racing to preempt each other with their launches. I don't know if that's what's happening in China or not, but it certainly had the flavor of like DeepSeek put their paper out and then a few hours later, here comes the Kimmy paper.
Their model isn't quite yet available. They said it will be available via API and presumably in their product soon, but it's not yet. So unclear what's going on in China. Did they intend to put these out on Trump's Inauguration Day or was that just an accident? It's kind of hard to believe that it is an accident. But at the same time, these folks are focused on unraveling the mysteries of AGI with curiosity. So maybe they don't care about when Trump is getting inaugurated. Maybe they're coordinating. I have a lot of questions about the dynamics that are going on in China behind this. But what is clear is that at least DeepSeek with this R1 model has joined the top tier of global AI developers. Potentially Moonshot with their Kimmy model could be there as well. But it's obviously hard to say that kind of stuff purely from the benchmarks. So we'll have to wait and get our hands on it before we can be too confident about that. What I want to do is walk through what stands out to me about this and try to make some sense of it. I suspect this will be the first of several conversations about this because this R1 story touches on so many different aspects of AI at the same time. The research itself is really important.
The consequences for just practical utility are significant as well. The gap between closed source and open source, also the gap between the West and China, I would say shrinking gaps at the moment. Certainly the gap between the West and China seems to have shrunk significantly from where it was a couple years ago. And that's just the start, right? Then there's, of course, the strategic dynamics, like why is China open sourcing this? What are they getting out of that? How, you know, if at all should the US respond? Does this challenge narratives that are increasingly dominant in the West about AI race? We've now seen no less than Alex Wang, the CEO of Scale AI, take out a full page ad in the newspaper calling the current situation an AI war, which I honestly totally hate and think is wildly irresponsible. You don't have to be a China dove to recognize that an AI war does not exist and would be bad for everyone. I hated to see that. We need to reconsider some of those framings in light of what we're seeing here. And certainly the strategy that we want to play, if our strategy is predicated on preventing China from doing certain things so that we can have certain advantage, so that we can solve certain problems, so we can be the good guys, I think that the window of opportunity that we have where we're gonna have this sort of unassailable AI lead looks quite short. And I think we'll understand that better as we go through the research and understand how simple a lot of the stuff is, driving a lot of these significant advances in reasoning capability. But at the end of all that, what sort of policy response, if any, makes sense to this? Does it still make sense to think about pre-training as being the real measure of model power or the standards by which a model would qualify for some sort of special process, special government review, special government notification? I think that is highly questionable in light of the power of the reasoning paradigm because significant gains over pre-training are showing up with presumably a lot less compute. Also, they're being distilled into much smaller models. It does seem like we have crossed a meaningful threshold this week where prior to this week, there was not really a good quality reasoning model that I could run on a local machine. Now I've got a wide range of things that are open-source that I can download that have been trained specifically as reasoners, which I can further modify including with more reinforcement learning. This is really one of the most important stories that has come to the public in AI in a while. I want to help make some sense of it. I thought we would start and I am doing a screen share this time around. So if you are listening to this, I think it will be fine. I will plan to basically read everything that's important. If you're a more visual person, you want to see stuff on screen and be able to read along, I'll have the screen share on YouTube as well. I am using a variety of AIs to help make sense of this. You'll see me tabbing back and forth and looking at various sources. But it should be fine in audio format if that's what you prefer. Let's start by talking about what the R1 model is and how they created it. There's a couple of different flavors and I think there's a couple of different big takeaways from this. First of all, they just recently came out, not too long ago, with their DeepSeek V3 model. This made headlines on its own for being a top-tier model that was made incredibly cheaply. Zvi, as always, has great coverage of this. He calls it the $6 million model. This is a small percentage of what Western AI leaders are understood to have spent to train their top-tier, frontier models. A lot of work has gone into the efficiency there.

87 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000685431704