**Swix** (0:00)
Hello, hello, it's Swix again. It's July 1st, 2023 Happy Q3 to those of you who celebrate. It is one day after we launched our blog post and conference, I guess, on the Rise of the AI Engineer. For those of you who only joined us for the podcast, the Latent Space newsletter actually preceded the podcast. And yesterday, we put out our first newsletter post in a long time. So if you want to just head over to Latent.Space, you can check out the post about why our thesis is about the AI engineer and why we think that this is going to be a growing category and why we're putting on our first conference in October in San Francisco. Join us if you can. CFP sponsorships and attendee slots are open as of yesterday. So today, it's basically the July 4th weekend. Most people are going to take Monday off and Tuesday is July 4th. So we figured we would do a double sort of podcast swap with some of our favorite AI podcasts that we love and enjoy and wanted to share with you again, especially highlighting some of the things that we liked about their podcasts and then they're doing the same for us on their feeds. So today we're featuring Nathan Labenz of the Cognitive Revolution Podcast. They started around the same time as us, but then they've just went way harder than us, done twice the number of episodes and have covered way more than us in terms of computer vision, healthcare, investing in tech, safety and policy, curators, influencers, as well as exceptional AI founders, some of whom we also hope to have on our podcast in some time in the future.
And the story that we've picked out or the episode that we picked out is a recent one, but I think is extremely important as a theme for 2023, which is TinyStories from Microsoft Research led by Ronen Eldan and Yuanjie Li.
Since they actually published TinyStories, they actually also published PhiOne, which is a large language model. It's at 1.3 billion parameters. It's only trained on 800, 800 hours. So that's about a couple thousand dollars of training. And it scores above 50% on human eval, which as we all know from the Replet episode that we did is an imperfect benchmark, but it is as far as people are concerned, the industry standard benchmark for code models. And it's basically performing at an equivalent level of models that are 10 times its size and 100 times its data set.
And so that's an interesting phenomenon that basically points to data quality as the new dimension for which model trainers are optimizing for. We've talked about various dimensions of scaling, from scaling the number of parameters to scaling the data set size, to scaling the amount of compute that you spend.
But this is the first time where it's a little bit harder to talk about this because there's no, it's really hard to quantify, but here we're scaling quality for the same amount of length or the same number of tokens that are invested in the training process. Yeah, but this interview is about TinyStories, which is the earlier work that informed the PHY1 model. And TinyStories is very endearing. It's only focused on the reading level of like a three to four year old, but it actually makes a lot of interesting points. The thing that model researchers pick up on is that it uses synthetic data sets that were generated by GPT-3 and 4 But for practitioners, I think the more useful and interesting or mind blowing thing is that it can generate very coherent stories from a tiny language models. Like I'm talking about less than 10 million parameters.
And they have this comparison, which I'm going to stick in the show notes, where they compare a story that was generated by TinyStories model, there's three million parameters, comparing it to GPT-2 XL, which is 1.5 billion parameters models. So it's beating a 500 times larger model because it's focused on this kind of domain.
And this has a lot of interesting implications on interpretability because it's a much smaller model. We can visualize what's going on in the weights rather than having a big mass where we don't know anything. But also just for the future of domain specific models. I also like the analogy that they make in the podcast interview, which is that this kind of training models how humans learn. Which is first you learn to speak like a child, and then you learn adult language, or to talk like an adult.
As opposed to how we trade large models today, which is we just throw it in a deep end and just expect it to learn all of common crawl and to speak like a professor, to speak like a 4chan troll, to speak like textbook authors.
87 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000618988028