Why Image Generation Needs More Than Bigger Models with Fatih Porikli artwork

Why Image Generation Needs More Than Bigger Models with Fatih Porikli

The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

August 12, 2026

Text-to-image models have become remarkably good at producing realistic images. But realism isn’t the same as correctness. Ask for several distinct people, a specific composition, or a high-resolution image generated locally, and today’s models still struggle in surprising ways.
Speakers: Sam Charrington, Fatih Porikli

Topics: Technology, News, Tech News

**Sam Charrington** (0:00)
Thanks so much to our friends at Qualcomm for their continued support and sponsorship of today's episode. Qualcomm AI Research is dedicated to advancing AI to make its core capabilities, perception, reasoning and action, ubiquitous across devices. Their work makes it possible for billions of users around the world to have AI enhanced experiences on devices powered by Qualcomm technologies.
To learn more about what Qualcomm is up to on the research front, visit twimlai.com/qualcomm.
Once a frontier research problem in computer vision, text to image generation has reached a point where almost anyone can ask an AI system for a picture and get something that looks remarkably good. But looking good and being correct are not the same thing. Ask for several different people and the model may generate variations of the same face. Ask for a specific composition, identity, or number of subjects, and it may ignore those details. Push towards higher resolution or local generation and quality speed and memory will quickly become constraints. The next frontier in image generation is closing this gap between plausible images and precise results. It's making these systems controllable, efficient, and reliable enough to consistently produce a high quality rendition of the image you actually ask for. One researcher at the forefront of this work is Fatih Porikli, Vice President of Technology at Qualcomm, whose team presented more than 20 papers at this year's CVPR, the Computer Vision and Pattern Recognition Conference. Here's Fatih explaining why image generation still has plenty of hard problems to solve.

**Fatih Porikli** (1:28)
Maybe we are also asking a single model to solve too many difficult problems at once. Think about what happens when you generate a scene with several people. Now, the model has to understand the problem, but I'm asking to model and then decide how many people should appear, determine where they should be placed, like the composition of the scene, reason about their interactions, because if there's a person, if there's another person, most likely there is some connection. Preserve the identity of the person. We can give, okay, this is my daughter, this is my son, and I want them to be in the picture, not like any random person, and finally render everything together in all a single process. So maybe this is too much.

**Sam Charrington** (2:15)
I'm Sam Charrington and this is the TWIML AI Podcast. For over a decade, I've been exploring the ideas and observations shaping the future of AI through conversations like this one that help you understand what's real, what's next, and what matters. Let's jump in.
Yeah, that's a big difference that I see between, I think even this conversation and the conversation we had last year, the performance and capability of text to image models has improved significantly. And not just performance and capability, but I think accessibility, like now ask whatever your favorite LLM is to generate an image, and it will do a really, really good job. And it does kind of beg this question of the computer vision community, like what's left to do? If this problem is solved to this degree, what's left for us to work on?

**Fatih Porikli** (3:19)
That's a fair question. T2I models, these text to image generation models, or image to image generation models have become incredibly good at producing very realistic images. The lighting looks natural, right? The details look right, and the overall quality can be amazing. As many things we do, there's initial excitement, there's great work coming up. But still, if you look into that one dive deeper, you realize there are many things to still be accomplished. One was controllability. Generating multiple people in the same image, people look almost identical, faces kind of blend together. And also, we showed in the past, we can run such models on user's devices. You do not need to rely on a cloud service provider. I think most of the models still are limited to 1K, 1K resolution. But then you want to go beyond that. That is the challenge, which has not been actually addressed before.

**Sam Charrington** (4:21)
Okay, so what I'm hearing in there is that we've made a lot of progress, but there's still work to be done.
And when you think about that work, some of the big buckets include controllability, the ability to really get the models to focus on the way you describe the task, or focus on the output that you want. And then you mentioned in there quality, so fewer artifacts, sharper images. And then you mentioned efficiency. We need to keep up with our ability to run the latest and greatest models on the device. So those are like three chunky buckets for researchers to continue to work in.

37 more minutes of transcript below

Thousands of transcripts fetched by people building searchable podcast archives

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/YOUR_EPISODE_ID