Data, data, everywhere - enough for AGI? artwork

Data, data, everywhere - enough for AGI?

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

April 13, 2024

In this podcast, Nathan and Nick dive deep into the data requirements for achieving Artificial General Intelligence. They explore the current paradigms, the role of data in approximating intelligence, and the scaling trends for GPT models.
Speakers: Nick Gannon, Nathan Labenz
**Nick Gannon** (0:00)
Oftentimes, people's conceptions of AI progress seem to be more so derived from aggregating the sentiments of the crowd than any core ground-up framework. This is something often I do as well, but we want to avoid reducing AI as a concept to an index that we're sort of longer, short, bearish, and bullish, overpriced, underpriced. Because doing so makes our models of the AI space sprout in other people's opinions of AI, rather than in any facts of the case. It's becoming awfully clear to me that these models are truly approximating their datasets to an incredible degree. What this manifests is trained on the same dataset for long enough, pretty much every model with enough weights in training time converges to the same point. Improvements in data quality and improvements in algorithmic architectures can be viewed as reducing the scale requirements to reach this human level performance in generality.
Across a large range of tasks.

**Nathan Labenz** (1:01)
Hello and welcome to The Cognitive Revolution, where we interview visionary researchers, entrepreneurs, and builders working on the frontier of artificial intelligence. Each week, we'll explore their revolutionary ideas, and together we'll build a picture of how AI technology will transform work, life, and society in the coming years. I'm Nathan Labenz joined by my co-host, Erik Torenberg. Hello and welcome back to The Cognitive Revolution. Today, we're diving deep into one of the most critical questions in AI. If we think that we can build some form of AGI by simply scaling up the current paradigm, is there enough high-quality data in the world to get us there? Joining me on this journey is Nick Gannon, master's degree data scientist by day, AI scout by night. He conducted extensive research to gather the numerous numbers that we'll be discussing over the next hour. We begin by extrapolating the trend set from GPT-2 to GPT-4, to set a budget for a hypothetical GPT-5. Then, we consider the total data volume generated across domains like email, social media, YouTube, genomics, and astronomy, attempting to determine just how much of humanity's raw data output would need to be high-quality to achieve this ultimate goal.
We also work backward from the scale of compute that we might expect to have in the future, asking how much data we'd need to use it all effectively.
As you'll hear, this episode is full of interesting numbers and useful comparisons. Our goal is to help you anchor your AI worldview to realistic ranges for the key AI inputs of data and compute, enabling you to better contextualize the growing volume of new research, data sets, and models that you'll have no choice but to process with increasing speed going forward.
While following the many order-of-magnitude calculations that we work through will likely require more focus than our typical episode, I personally believe the extra effort is worth it. In a world where the White House has set 10 to the 26th flops as the threshold for reporting training runs, I think anyone seriously tracking AI progress should actively build intuition around key reference numbers.
As always, if you value this work, we appreciate it when you share it with friends. And for this experimental episode in particular, we especially request your feedback, either via our website, by DMing me on your favorite social network, or if you loved it, via a review on Apple Podcasts or Spotify. I don't see anyone else doing quite this kind of work, but is that for good reason? Or would you like to see us invest more in these high-level guides? In any case, your input will shape our future decisions.
Now, without further ado, here's my discussion about the scale requirements for AI training data with Nick Gannon. Nick Gannon, welcome to The Cognitive Revolution. Thank you. So we are here to talk about data. And I've been really intrigued by some of the analysis that you've brought to bear on this question of what data exists, what's out there. Are we going to run out of it? What are the different modalities? So this is a scouting report episode that really just tries to get our arms around the scope, the scale, and the nature of data to the best of our ability. And I really appreciate all the work that you've put in to try to answer these questions. I'm excited to learn a lot from your analysis.

**Nick Gannon** (4:17)
Yeah, absolutely. Yeah, thanks for letting me share. So diving in here, the general premises, what are the data requirements to brute force your way to systems that are as generally intelligent as humans across essentially all cognitive tasks? Oftentimes, people's conceptions of AI progress seem to be more so derived from aggregating the sentiments of the crowd than any core ground up framework. This is something often I do as well, but we want to avoid reducing AI as a concept to an index that we're sort of longer, short, bearish, and bullish, overpriced, underpriced. Because doing so makes our models of the AI space grounded in other people's opinions of AI, rather than in any facts of the case. So instead of this AI perspective of crowd sentiment analysis aggregation, it makes more sense to live within an explanatory paradigm to explain the current state of affairs. In AI, the current paradigm that sits at the core of the hyper-scales being OpenAI, DeepMind, and Enthropic is the scaling hypothesis of intelligence. And Ilya puts it pretty well outlining two fairly simple premises for the scaling hypothesis that sort of underlies the AI strategy of these three firms. Premise one being really just this if brain bigger than brain smarter premise.

40 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000652333924