Efficiency is Coming: 3000x Faster, Cheaper, Better AI Inference from Hardware Improvements, Quantization, and Synthetic Data Distillation artwork

Efficiency is Coming: 3000x Faster, Cheaper, Better AI Inference from Hardware Improvements, Quantization, and Synthetic Data Distillation

Latent Space: The AI Engineer Podcast

September 3, 2024

AI Engineering is expanding! Join the first 🇬🇧 AI Engineer London meetup in Sept and get in touch for sponsoring the second 🗽 AI Engineer Summit in NYC this Dec! The commoditization of intelligence takes on a few dimensions: * Time to Open Model Equivalent: 15 months between GPT-4 and Llama 3.
Speakers: Alessio, Swyx, Nyla Worker
**SPEAKER_2** (0:35)
So, what we see in Latent Space is the importance of efficiency in all forms, from sample efficiency for spending limited training compute on limited data, and increasingly towards inference efficiency for increasingly demanding use cases like local LLMs, real-time AI NPCs and edge AI. However, we've never really developed any intuition for the trends in efficiency over time. For example, from 2020 to 2023, the price of GPT-3-level intelligence dropped from $60 per million tokens to $0.27 with the mixtral price war of December 2023 See show notes for charts and data. As for GPT-4-level intelligence, it took just over a year for GPT-4 to be matched by Llama 370B and GPT-4 Turbo to be beaten by Llama 3405B in open source, causing blended cost per million tokens to freefall from over $30 for Claude 3 Opus and the original GPT-4, down to under $3 for Llama 3405B. Of course, OpenAI themselves have not stood still, slashing the price of GPT-40 by 30 times with GPT-40 Mini. Yes, you heard that right. GPT-40 Mini is 3.5% the price of GPT-40, yet ties with GPT-40 Turbo on LM Sys. When the price of intelligence is falling by over 90% every year, what are the driving forces and how should AI engineers plan for this? It turns out that this has happened before in computer vision, which has in theory almost 3,000 times throughput improvement over the last six years while holding quality constant due to stackable improvements in efficient GPUs, quantization, pruning and distillation. We invited Nyla Worker of Nvidia and convai, and most recently our gracious track host of the GPUs and inference track at the AI engineer World's fair, who first made this observation to Swix to help talk us through the past, present and future use cases of efficient AI inference. Note that this was recorded before Nyla joined Google AI to work on efficiency. So you can expect more great efficiency work coming from her on the Gemini team. While at convai, Nyla also worked on the famous ramen demo with Nvidia that won multiple Best of CES awards from PC Gamer, Tom's Guide and the Shortcut earlier this year. So we also chatted about this year's CES, Computex and the future of AI and PCs. In Latent Space News, look out for our upcoming London and NYC meetups on the community page and of course feel free to start your own and simply let us know. Watch out and take care.

**Alessio** (3:34)
Hey everyone, welcome to the Latent Space Podcast. This is Alessio, partner and CTO and resident at Decibel Partners, and I'm joined by my co-host Swix, founder of Small AI.

**Swyx** (3:43)
Hey, and today we are in the remote studio with Nyla Worker. Welcome Nyla. Good to see you.

**Nyla Worker** (3:48)
Good to see you all.

**Swyx** (3:50)
So we try to introduce people based on their professional profile and then let you fill in the blanks. So you did astrophysics research at Carlton College, and then you made your way into machine learning. We're going to talk about your time at eBay. But most recently, you spent four years at Nvidia working on everything from synthetic data to cloud container offerings, and now currently you're director of product management at convai. What should people know about you that maybe it's not super obvious on your LinkedIn that it encapsulates your life journey so far?

**Nyla Worker** (4:20)
Yeah, I think the thing that is not very obvious is the transition from astrophysics research to AI and how that happens. So within astrophysics, what I was doing on my freshman year of college was categorizing whether this was a supernova Rembrandt or like an exoplanet. And while that sounds all cool and incredible, it's literally looking at images of like oxygen and sulfur and selecting manually each region. And it is extremely boring, as I shall I say. So I then found a paper from 1996 called Source Extractor, or like he called it Sextractor for some reason. And it was a multi-layer perception network that had been trained on synthetic data to categorize whether this was a star or a galaxy. That led me to see that there was this massive optimization machine that when fed with right data, it could perform an automate task such as this kind of manual classification. That made me want to learn, how do you train these things? How do you deploy them effectively? And if it's useful for just classifying galaxies, what other applications are there out there where we show a bunch of data and just train these functions to just predict the next word? In the case of LLMs or predict what is, is this a cat or a dog and things like that. So then I went to computer vision research, particularly scaling the training of deep neural networks. Back then I was using CPUs, doing it wrongly, of course.

50 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000668186527