**Conviction** (0:05)
Hi, listeners, and welcome to another episode of No Priors. Today, we're talking to Suhail Doshi, the founder of Playground AI, an image generator and editor. They've been open sourcing foundation diffusion models, most recently Playground V25. We're so excited to have Suhail on to talk about building this model in conjunction with the Playground community and the future of AI pixel generation. Welcome, Suhail.
So this is your third company, you started Mixpanel, Mighty, now you're working on Playground. How did you decide this was the next thing?
**Suhail Doshi** (0:36)
I think like back in April of 2022, I think that was just a place that time, it was like GPT-3, 5 kind of came out and then Dolly 2 came out.
And I was actually working on the second company, Mighty. And at that time, I was trying to figure out how to like do something with AI inside of a browser address bar. But when I saw Dolly 2 came out, it was just this like very big, strange eye-opening moment where I think a lot of people didn't think that we'd be able to do like weird, interesting art things so soon. And so, and I think then soon after that, I think Stable Diffusion came out around June or July of that same year and I got early access, maybe a couple of weeks access to early to SD14.
And I just kind of blew my mind what people could do with that. And I just thought that it seemed odd that all of this was being done in a Google Colab Notebook.
Shouldn't there be like a UI that makes it really easy? That sort of thing.
**Conviction** (1:34)
From the start, were you just thinking we will open source, we will train our own models from scratch? Did you think about other modalities?
**Suhail Doshi** (1:43)
Yeah, I mean, there have been a lot of people that thought I should do like something in music, but I just, because music has been like a huge hobby of mine for like six years or so. I like produce music, but I just couldn't wrap my brain around like what useful thing I would end up making for people. Although now there's like a lot of very interesting, cool, useful things for music.
And then it seemed like a lot of people were very focused on language and I had really enjoyed, I already work with lots of creative tools. Like when I was in high school, I used to make logos or I would make music or whatever.
So I was excited that finally I could find something where it was a combination of creativity, tooling. Images have really amazing built-in distribution. People wanna share those kinds of things. So it ended up just being this perfect thing that I was excited to work on.
**Conviction** (2:32)
How do you think the overall landscape for competition is different in language versus images versus music? How did you think about in what ways you guys would want to build advantage and stand out?
**Suhail Doshi** (2:48)
I think with language, there's like, I don't know, I don't know how many language companies there are. You guys would probably know better than me, but it seems like there's like over 20 And then maybe like five or like eight of them have a billion dollars worth of funding. I also didn't wanna work on something if there were already extremely passionate people really working hard at that thing. People that I like really respected that were working on that thing. And so at the time with images, there was just sort of, I think there was Mid Journey, there was OpenAI doing some Dali stuff, and then you saw sort of Stable Diffusion.
But for some of these companies, it didn't seem like there was going to be a long standing concerted effort to keep making them better. It was sort of unclear who was doing this as a fun demo versus who was doing this as something they would spend and invest tons of their time in. And so once I had kind of figured out to what extent OpenAI was gonna invest in it, or to what extent seemed like the folks at Stability AI were sort of focused on seven different kinds of things. And I just thought, hmm, there's just not enough people that wanna do this one thing and do it really, really great. So I think for me, it was just about were there enough capable people that wanted to do this? Can you talk a little bit about the specific direction you decided to take with Playground as well? And I know you thought really deeply about some of the applications or use cases for it. So I was just curious if you could share a bit more about that.
21 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000652858596