The good, the bad, and the AI apps artwork

The good, the bad, and the AI apps

The Stack Overflow Podcast

July 3, 2026

Ryan welcomes Benny Chen, co-founder of Fireworks AI, to the show to explore what actually makes an AI application good or not, how to balance qualitative signals with quantitative metrics when evaluating AI, and how open-source eval protocols and community efforts are setting the standard for AI...
Speakers: Ryan Donovan, Benny Chen
**Ryan Donovan** (0:07)
Hello, and welcome to The Stack Overflow Podcast, a place to talk all things software and technology. I am your host, Ryan Donovan, and today we're talking about why it's so hard to get the definition of good right for AI, and how to get that right at scale. So my guest for that is Benny Chen, who's the co-founder of Fireworks AI.
So welcome to the show, Benny.

**Benny Chen** (0:30)
Thanks for having me.

**Ryan Donovan** (0:31)
Before we get into our topic today, we like to get to know our guests. How did you get into software and technology?

**Benny Chen** (0:37)
In middle school, my mom gave me a really, really old computer. It couldn't really run Windows reliably. So I don't know how I came across it. But I started trying to use Fedora. But people forget back then even to install FFmpeg, to run MPlayer was a big hassle.
So I had to learn how to do scripting. And as you know, FFmpeg's command line is basically a programming language. So in order to watch anime, I started learning how to do programming.
20 years ago, it was a really serious thing to install a codec. Codec to run was a feat.

**Ryan Donovan** (1:17)
You know, I hear a lot of people who got in through video games, you might be the first who got in through anime.
Then you came to found an AI company. How did that come about?

**Benny Chen** (1:26)
Yeah, I think it was around 22 I was responsible for ads capacity planning for meta to a certain extent, and we were looking at all the projections, bottom to top down, doesn't matter how you look at it. Demand for AI workflow was going through the roof. In retrospect, that's nothing compared to what we're looking at today. And I remember it was like, oh, we had to hit like maybe, I think half a million cards for a certain program to make sense.
These days, the numbers are child's play, compared to what Elon and Zack is doing. And I remember in 22, what motivated me a lot was, hey, the workload is going through a roof, and for us taking off, and we're talking about hundreds of megawatts. Even one day, there will be gigawatts. And now, everyone is waving their hand and be like, yes, we're building a few data centers as an operational thing. I worked on supporting the Meta ASIC program with Intel in 2017
So in 2017, the Meta leadership was looking at Google's TPU program and be like, we need something similar. We were working with Intel together on that. I got the feedback that I was one of those annoying ads people that was trying to push it through in 12 months, and it was breaking a lot of systems. And now you see Elon standing up colossus in three months.
The writing was on the wall that AI infrastructure was taking off, and then it was a good time in 22 to do this. But we did not expect the genitive AI workload to take off like this. This is insane.

**Ryan Donovan** (3:08)
It has taken off at a pretty insane to where I think everybody in software development has to touch it at some point. But I think one of the things I want to talk to you about, something I've been thinking about too, is how do you define what good is for an AI application?
I think this has been a problem in software engineering in general, but it is more so with a non-deterministic process with AI. So how do you think about defining good for an AI workload or an AI application?

**Benny Chen** (3:41)
Articulating what is good is very, very difficult.
At the end of the day, we want to make sure our customer can make their customer happy.
We are primarily a B2B SaaS company. We support a lot of AI applications. They have a very difficult time competing with OpenAentropic. So for them, at the end of the day, they look at their customer satisfaction metrics. They look at all the very sparse outputs that their customers provide to them. We need to provide the infrastructure and the platform to help people turn those really sparse signals into something more measurable, something more dense that people can use on a daily basis, as well as using it for model fine-tuning or using it to just pick open source models as well. So what is good often translates into something that is repeatable, that's something you can iterate on on a daily basis. That's also more dense instead of just sparse words that people get on, like, I don't know, app review or like thumbs up or things like that.

18 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000775272820