Image Generation and Visual Intelligence with Black Forest Labs artwork

Image Generation and Visual Intelligence with Black Forest Labs

Practical AI

July 2, 2026

How has AI image generation evolved from blurry outputs to powerful visual intelligence models?
Speakers: Daniel Whitenack, Chris Benson, Dustin Podell
**SPEAKER_1** (0:02)
Welcome to the Practical AI Podcast, where we break down the real world applications of artificial intelligence and how it's shaping the way we live, work, and create. Our goal is to help make AI technology practical, productive, and accessible to everyone. Whether you're a developer, business leader, or just curious about the tech behind the buzz, you're in the right place.
Be sure to connect with us on LinkedIn, X, or Bluesky, to stay up to date with episode drops, behind the scenes content, and AI insights. You can learn more at PracticalAI.fm. Now, on to the show.

**Daniel Whitenack** (0:41)
Welcome to another episode of the Practical AI podcast. This is Daniel Whitenack. I am the CEO at Prediction Guard, and I'm joined as always by my cohost, Chris Benson, who is a principal AI and autonomy research engineer. How are you doing, Chris?

**Chris Benson** (0:56)
Hey, I'm doing great. I can't wait to get into today's conversation. It's going to be fun.

**Daniel Whitenack** (1:02)
Yes.
For an audio podcast, we're going to talk about a lot of interesting visual things. Maybe before we get started, just a little teaser, Practical AI is posting some videos on YouTube now. If you do consume podcasts that way, you might go check us out on our YouTube page. But speaking of images, videos, and more specifically, image generation, really excited to have with us today, Dustin Podell, who is co-founder and researcher at Black Forest Labs. Welcome, Dustin.

**Dustin Podell** (1:38)
Yeah. Thanks for having me, guys. It's really great to be here.

**Daniel Whitenack** (1:41)
Yeah. I know that Black Forest Labs does more than just raw image generation. There's a lot of workflow related things, hardware optimization, all sorts of cool stuff you're involved with. But as we get into some of that, I'm wondering if you can just help our audience with a bit of a state of image generation methods and workflows for the industry. We've talked on the show before about diffusion models, and we'll link some of those episodes in the show notes maybe.
But a lot has happened, right? There's a lot of people working on a lot of interesting things, and I'd love to understand over the last year, what are some of those main points that might be good for people to orient themselves to where things are at now?

**Dustin Podell** (2:28)
Yeah, no, it's a good question. I mean, if I'm allowed to take it even a little bit further back, I mean, the state of-

**Daniel Whitenack** (2:34)
Absolutely.

**Dustin Podell** (2:35)
Yeah, the state of image gen, video gen, generative models as a whole has gone crazy, so to speak, in the last three or four years.

**Daniel Whitenack** (2:51)
Where are we now?

**Dustin Podell** (2:52)
Where we came from about four years ago, where I first entered into the scene, so to speak, is we were at models that were essentially just doing little blobs of color that were related a little bit to where you were with the prompt, okay, a lighthouse on the beach, this and that.
You would get something that, okay, that vaguely looks like it and you would show someone and it'd be, yeah, I can see it, I guess, and maybe it would interest a few nerdy people. Then now today, we're at the point where, if you've probably been seeing plenty of this stuff online, we're seeing whole short films made entirely with AI generation, where certain scenes are almost entirely indistinguishable from reality, so to speak.
I would say we've come quite far, but I will also say the core of the technology hasn't actually changed that much. It's been a pretty nice forward progress. I don't want to dive immediately into anything technical here, but for anyone who's been paying attention, I'm sure, or anyone who really hasn't been paying attention, this probably came a bit out of nowhere, so.

**Chris Benson** (4:00)
I was going to say, I guess with the broadly attention by the general public is so much on more of the LLM generative world in terms of that stuff, and everyone's finally on apps regardless of whether they're technical or not. And so I think a lot of people kind of miss the tremendous advancements you guys are making on that side. And so, like, you know, could you take a quick moment, and now that you've kind of done the highest level, maybe step through some of the things, and people may remember, like we talked about stable diffusion and things, a couple of those things that have kind of led up to what we're going to dive into today with some more specifics. And as Daniel mentioned, we can offer some links for past conversations if people want to dive into those specifically. But that would really be interesting to kind of hear distinct points on that timeline as we hit to this point.

42 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000775138797