Owning the AI Pareto Frontier — Jeff Dean artwork

Owning the AI Pareto Frontier — Jeff Dean

Latent Space: The AI Engineer Podcast

February 12, 2026

From rewriting Google’s search stack in the early 2000s to reviving sparse trillion-parameter models and co-designing TPUs with frontier ML research, Jeff Dean has quietly shaped nearly every layer of the modern AI stack.
Speakers: Alessio Fanelli, Shawn Wang, Jeff Dean
**Alessio Fanelli** (0:04)
Hey everyone, welcome to the Latent Space Podcast. This is Alessio, founder of Kernel Labs, and I'm joined by Swyx, editor of Latent Space.

**Shawn Wang** (0:10)
Hello, hello. We're here in the studio with Jeff Dean, chief AI scientist at Google. Welcome.

**Jeff Dean** (0:14)
Thanks for having me.

**Shawn Wang** (0:16)
It's a bit surreal to have you in the studio. I've watched so many of your talks, and obviously your career has been super legendary. So I think the first thing must be said, congrats on owning the Pareto Frontier.

**Jeff Dean** (0:30)
Thank you. Pareto Frontiers are good, and it's good to be out there.

**Shawn Wang** (0:34)
Yeah. I think it's a combination of both.
To own the Pareto Frontier, you have to have frontier capability, but also efficiency, and then offer that range of models that people like to use. Some part of this was started because of your hardware work, some part of that is your model work. And I'm sure there's lots of secret sauce that you guys have worked on accumulatively. But it's really impressive to see it all come together in this steadily advancing frontier.

**Jeff Dean** (1:05)
Yeah. I think as you say, it's not just one thing, it's like a whole bunch of things up and down the stack. And all of those really combine to help make UNOS able to make highly capable large models, as well as software techniques to get those large model capabilities into much smaller, lighter weight models that are much more cost effective and lower latency, but still quite capable for their size.

**Alessio Fanelli** (1:31)
Yeah. How much pressure do you have on having the lower bound of the Pareto Frontier too? I think the new labs are always trying to push the top performance frontier because they need to raise more money and all of that. And you guys have billions of users. And I think initially when you worked on the CPU, you were thinking about if everybody that used Google we used the voice model for like three minutes a day, they were like, you need to double your CPU number. Like, what's that discussion today at Google? Like, how do you prioritize frontier versus like we actually need to deploy it if we build it?

**Jeff Dean** (2:03)
Yeah. I mean, I think we always want to have models that are at the frontier or pushing the frontier because I think that's where you see what capabilities now exist that didn't exist at the sort of slightly less capable last year's version or last six months ago version.
At the same time, you know, we know those are going to be really useful for a bunch of use cases, but they're going to be a bit slower and a bit more expensive than people might like for a bunch of other broader use cases. So I think what we want to do is always have kind of a highly capable sort of affordable model that enables a whole bunch of, you know, lower latency use cases. People can use them for agentic coding much more readily. And then have the high end, you know, frontier model that is really useful for, you know, deep reasoning, you know, solving really complicated math problems, those kinds of things. And it's not that one or the other is useful. They're both useful. So I think we like to do both. And also, you know, through distillation, which is a key technique for making the smaller models more capable, you know, you have to have the frontier model in order to then distill it into your smaller model. So it's not like an either or choice. You sort of need that in order to actually get a highly capable, more modest size model.

**Alessio Fanelli** (3:24)
Yeah. I mean, you and Jeffrey came up with distillation in 2014

**Jeff Dean** (3:28)
Don't forget, L'Oreal Vignoles as well.

**Alessio Fanelli** (3:31)
A long time ago, like, I'm curious how you think about the cycle of these ideas, even like, you know, sparse models. And, you know, how do you re-evaluate them? How do you think about in the next generation of model, what is worth re-visiting? Like, yeah, they're just kind of like, you know, you work on so many ideas that end up being influential, but like, in the moment, they might not feel that way necessarily.

**Jeff Dean** (3:52)
Yeah, I mean, I think Distillation was originally motivated because we were seeing that we had a very large image data set at the time, you know, 300 million images that we could train on with, you know, I forget, like 20,000 categories or something, so much bigger than ImageNet. And we were seeing that if you create specialists for different subsets of those image categories, you know, this one's going to be really good at sort of mammals, and this one's going to be really good at sort of indoor room scenes or whatever. And you can cluster those categories and train on an enriched stream of data after you do pre-training on a much broader set of images, you get much better performance if you then treat that whole set of maybe 50 models you've trained as a large ensemble. But that's not a very practical thing to serve, right? So distillation really came about from the idea of, okay, what if we want to actually serve that and train all these independent sort of expert models and then squish it into something that actually fits in a form factor that you can actually serve. And that's not that different from what we're doing today. Often today, we're, instead of having an ensemble of 50 models, we're having a much larger scale model that we then distill into a much smaller scale model.

71 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000749498954