Why Diffusion Will Win AI Inference with Inception Co-Founder and CEO Stefano Ermon artwork

Why Diffusion Will Win AI Inference with Inception Co-Founder and CEO Stefano Ermon

No Priors: Artificial Intelligence | Technology | Startups

September 18, 2026

As generative AI hits hardware and latency bottlenecks, Stanford professor, diffusion pioneer, and Inception co-founder and CEO Stefano Ermon is betting on a radical new architecture.

Speakers Sarah Guo, Stefano Ermon

TopicsTechnologyBusinessEntrepreneurship

Sarah Guo (0:05)

Hi, listeners. Welcome back to No Priors. Today, I'm here with Stefano Ermon, who is a longtime Stanford professor and now co-founder and CEO of Inception. Stefano has a extraordinarily broad body of work around generative modeling, but is especially well-known as one of the fathers of diffusion. We talk about his company challenging the large labs, and why speed and efficiency are going to be the name of the game in AI over the next few years. Welcome, Stefano.

Stefano, thanks so much for being here.

Stefano Ermon (0:35)

Great to be here.

Sarah Guo (0:35)

I would love for us to just start with a little bit of your research background and how you ended up starting your company.

Stefano Ermon (0:42)

For sure. I've been doing research in generative models for basically my entire career. I started at Stanford in 2014 as an assistant professor, and I was working on building generative models. Back then, the research area was not particularly hot.

The models were not quite working well. We were still building little generative models over MNIST, and it was a big success if you could generate these grainy images of digits. It was even hard to publish papers back then on that topic, and you had to justify training a generative model as a way to learn features from unlabeled data that then could maybe help you better at supervised learning, because that was the thing that everybody cared about. But then things took over, of course, and so I was at the right place at the right time, working on the right thing, and so I've been doing research in that space since the beginning, basically.

Sarah Guo (1:40)

Did you have, besides curiosity in the area, a personal hope for what the models would do back in 2014 and 15?

Stefano Ermon (1:48)

Yeah, I mean, I always felt like that was going to be the, that was the right way to think about kind of like learning from unlabeled data, that like building a generative model is really the right way to make sure you understand the structure in the data. That was kind of like the way I was getting at.

I was not even dreaming about the kind of capabilities that these LLMs that we have today could do that. But I was thinking more from a world models perspective. I was working a lot on images and so thinking about, okay, I have a world model. I can imagine what's going to happen if I were to stand up and walk out the door. I can kind of picture that in my mind and that's important to make decisions and kind of model predictive control when having this kind of model of the world requires some generative capabilities. So I always felt like that's the right direction to work on. I felt like this is going to be very hard as a problem. It's going to keep me busy for my whole career and it's a good problem to work on. Then of course I was very wrong and things evolved much faster than I was expecting.

Sarah Guo (2:51)

Yeah, I think that's kind of universally true though. Walk me through the state of your research and how that led you to start the company.

Stefano Ermon (2:59)

Yeah, so I was working on generative models of images. I'm initially working on autoregressive models, which were very slow and kind of very blurry. And then VAEs and then GANs took over.

Sarah Guo (3:11)

Yes.

Stefano Ermon (3:11)

And back then, we were very unhappy with the state of generative models for images. The GANs, they worked, but they were very unstable to train, very hard to reproduce results. And so we were trying to see, is there a way to build something that is as good as a GAN, but it's more principled? And so we started working on score base, the generative models, which are basically what eventually became diffusion models back in 2019 with my PhD student, Young Song. And so we kind of came up with this idea of, let's train an oral network to denoise images. And if you can denoise an image, then you really are understanding enough about the structure of the image, that it should be possible to build a generative procedure based on these denoisers. And that basically became the kind of underlying technology of modern diffusion models, where instead of generating images, left to right, one pixel at a time, you kind of like start from pure noise and then you gradually refine the object until you get a clean picture at the end.

And that started back in 2019 in my lab and then it kind of took over the space. And even today, the best models for image generation, video generation, music to some extent, a lot of the protein stuff, they are based on diffusion. And so my group has worked a lot on various kinds of diffusion models, technique for accelerating them, to generate samples very quickly, to improve the quality of these models. And so since we were able to get them to work on images, I started thinking about, how do we get diffusion models to work on text or code generation and discrete objects? Like, is there a way to move beyond autoregressive models to something that it's more parallel, more with built in error correction? And so I've been doing a bunch of research at Stanford on getting diffusion models to work on text and code generation.

29 more minutes of transcript below

Thousands of transcripts fetched by people building searchable podcast archives

Fetch the whole transcript

The demo key returns a sample episode in full, no card needed:

request
curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Markdown with the speakers named, for your notes, your knowledge base, or anything that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.

Using your own key:

request
curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000790491295