Distributed Training, Decentralized AI: Prime Intellect's Master Plan to Make AI Too Cheap to Meter artwork

Distributed Training, Decentralized AI: Prime Intellect's Master Plan to Make AI Too Cheap to Meter

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

February 5, 2025

Vincent Weisser and Johannes Hagemann, founders of Prime Intellect, join a conversation on the Cognitive Revolution to delve into distributed training, decentralized AI, and their vision for a future where compute and intelligence are widely accessible.
Speakers: Vincent Weisser, Nathan Labenz, Johannes Hagemann
**Vincent Weisser** (0:00)
Execution is cheap. Ideas are worth everything, right? In a world where you can just like, it almost inverts to the current reality. And I think it would just lead to like billions of startups. We don't buy like hundreds of billions of computers. And that sounds like we're not a hotel, we're more like Airbnb or like we're more marketplace sitting on top even of other marketplaces.

**Nathan Labenz** (0:19)
Why does this matter? Multiple reasons, right? It's like in the limit, you know, it could create a sort of truly decentralized AI infrastructure that nobody can control.

**Johannes Hagemann** (0:29)
You obviously have a lot of coming up in terms of like improving all that algorithm, right? I think what we've done so far is just what we've realized those pseudo gradients were actually sent after those hundreds of steps. So it's not the actual gradients of the model, it's a difference between the beginning of the weights and the end state of the weights, after all those in-step updates.

**Vincent Weisser** (0:48)
In the intelligence age almost, you want to own a piece of a super intelligent system that is able to generate value, where you have actually had like access, like through your ownership in it, to the compute, to the intelligence.

**Nathan Labenz** (1:03)
Hello, and welcome back to The Cognitive Revolution. Today, I'm excited to share my conversation with Vincent Weisser and Johannes Hagemann, founders of Prime Intellect, whose mission is to make intelligence too cheap to meter by building foundational technology to support decentralized, collectively owned AI. Vincent and Johannes stand out for offering a positive vision of a future in which a wide range of AIs empower everyone simultaneously, amplifying each individual's abilities and improving societal resilience, while all actors implicitly check and balance one another's power. At the same time, they've articulated an ambitious master plan and shipped a number of notable milestone projects in pursuit of this goal. Part one of their plan is to build an international market for compute. And as of this writing, you can rent an H200 for $1.49 an hour via their website, primeintellect.ai. Part two is to build software frameworks for distributed training. And in late November, they released Intellect One, proving that distributed training can scale up to at least the 10 billion parameter level. Part three is to train high-impact science models. And the Metagen One model, developed in collaboration with the Nucleic Acid Observatory and others, and designed to be useful for pandemic detection, but architecturally incapable of generating new pathogens, is one of the best examples of a defense-favoring AI project that I've seen anywhere. Part four is to launch a decentralized protocol for collective ownership of AI models, and to collaboratively build towards aligned AGI that benefits all of humanity. While that still remains in front of them to do, given their track record to date, I would not bet against them making a meaningful contribution. We spent much of the first half of this conversation unpacking their vision. To be honest, I'm still not sure how realistic it is to expect that we can maintain a stable societal equilibrium with AI changing everything everywhere all at once. But then again, to be real, this is happening very fast, and I don't think anybody has articulated a credible big-picture plan so far. If that's true, and we're mostly just going to keep developing this technology as fast as possible and hope that the resulting AIs end up being mostly harmless by default, I do find a lot to like in their vision for a more decentralized and hopefully resilient balance of power, as opposed to a world dominated by a few major AI players. In the second half of our conversation, we get into the technical details, including both the fundamental challenges and recent progress in distributed training. Johannes walks us through the three main parallelization strategies used in model training, data, pipeline, and tensor parallelism, and discusses strategies like DeepMind's Deloco, which reduces communication overhead by allowing training nodes to process hundreds of steps before needing to aggregate gradients and sync model states. That they've managed to use this and a number of other optimizations to train a 10 billion parameter model across a globally distributed network of compute resources is impressive, and the latest from Google called Streaming Deloco, which was released just after we recorded, suggests that they have not yet hit any fundamental limits. It might still be very difficult to aggregate enough compute to train foundation models from scratch in a distributed fashion, but the recent shift toward reinforcement learning, which is far more inference-heavy and thus friendlier to distributed training approaches, strongly suggests that any number of network groups can probably muster enough compute to train whatever models they might like to reinforce into existence. Or, as Anthropix Jack Clark put it in a recent edition of Import AI, we might soon live in, quote, a world of models trained continuously in the invisible global compute sea. That world would almost certainly be simultaneously weird and beautiful and scary, but barring extremely draconian measures, the likes of which are well outside the Overton window today, it seems like the sort of thing that technologists can and will create, and that governments will have a very difficult time preventing or controlling. Perhaps in the end, we can only hope that DEAC, the Strategy of Accelerating the Differential Development of Decentralized Defenses, as exemplified by the MetaGene 1 model, will ultimately win out. As always, if you're finding value in the show, we'd appreciate it if you'd share it with friends, read a review on Apple Podcasts or Spotify, or leave us a comment on YouTube. We always welcome your feedback and suggestions via our website, cognitiverevolution.ai, or by DMing me on your favorite social network. For now, I hope you enjoy this discussion about a positive vision for, and the technical underpinnings of, decentralized AI development with Vincent Weisser and Johannes Hagemann of Prime Intellect.

124 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000689351188