**Alessio** (0:00)
Happy New Year, friends! Thanks for all the love on the Latent Space Live and 100th episode End of Year Recap. Your support has boosted us 30 places in the podcast charts, and that always helps us book great guests and organize more industry events for you. We don't say this enough, but thank you to everyone who has left a review on Apple Podcasts or subscribed to our new YouTube channel. Last year, we broke new ground when we interviewed our first public company CEO with Drew Houston and first technology cabinet member with Minister Josephine Teo, and first year with full coverage of leading labs across Meta, OpenAI, Anthropic, Raker, Liquid, and Google DeepMind. For our 101st episode, we are proud to introduce another first with our first anonymous guest. As Swyx mentions in the episode, Latent Space was started in the immediate aftermath of Stable Diffusion, and the uncredentialed software engineers it enabled, set the stage for the LLM wave that was to come with ChatGPT. The earliest winner of the Stable Diffusion tooling wars was SDWebUI, a Gradio app by the anonymous young creator Automatic1111, that quickly amassed over 100,000 Github stars for how it rapidly shipped plug-ins and usable interfaces for the rapidly growing Stable Diffusion ecosystem.
However, these days, the power tool of choice is now ComfyUI by today's guest ComfyAnonymous, who is gracing us with his first ever podcast appearance today. The shift from Automatic1111 to ComfyUI reflects a shift away in the image diffusion space from prompting and tweaking settings in 2022 to more complex and parallel workflows chaining together different models and orchestrating long-running operations that can also include video processing, visualized on an intuitive canvas instead of long YAML or code blocks. Because ComfyUI is open-source, there are now multiple Y Combinator startups built off of a Comfy workflow, or offering ComfyUI as a service directly. Interestingly enough, this same workflow tooling has not seemed to take off for other modalities yet, but perhaps 2025 is the year diffusion tooling diffuses to non-image domains. In other news, we have just announced the second AI Engineer Summit in New York City. We are bringing back the surprisingly successful AI Leadership Track from World's Fair. And also the single-track AI Engineering Track is now wholly focused on agents at work. If you are building agents in 2025, this is the single best conference to attend. Head to apply.ai.engineer and see you there. Watch out and take care.
**Swyx** (2:55)
Hey everyone, welcome to the Latent Space Podcast. This is Alessio, partner and CTO at Decibel Partners, and I'm joined by my co-host Swyx, founder of Small AI.
**SPEAKER_3** (3:03)
Hey everyone, we are in the Chroma Studio again, but with our first ever anonymous guest, Comfy Anonymous. Welcome.
**SPEAKER_4** (3:11)
Hello.
**SPEAKER_3** (3:12)
I feel like that's your full name. You just go by Comfy, right?
**SPEAKER_4** (3:15)
Yeah, well, a lot of people just call me Comfy, even though, even when they know my real name. Hey, Comfy.
**Swyx** (3:23)
Swyx is the same. Not a lot of people call you Shawn.
**SPEAKER_3** (3:27)
Yeah, you have a professional name, right? People know you by and then you have a legal name. Yeah, that's fine. How do I phrase this?
People who are in the know know that Comfy is the tool for image generation and now other multimodality stuff. I would say that when I first got started with Stable Diffusion, the star of the show was Automatic 111, right? And I actually looked back at my notes from 2022-ish. Comfy was already getting started back then, but it was kind of like the up-and-cover and your main feature was the flowchart. Can you just kind of rewind to that moment, that year, how you looked at the landscape there and decided to start Comfy?
**SPEAKER_4** (4:02)
Yeah, I discovered Stable Diffusion in 2022, in October 2022 And well, I kind of started playing around with it. Yes, I did. And back then, I was using Automatic, which was what everyone was using back then. And so I started with that because when I started, I had no idea how diffusion models work, how any of this works.
**SPEAKER_3** (4:27)
Oh yeah. What was your prior background as an engineer?
**SPEAKER_4** (4:30)
Just a software engineer. Yeah, boring software engineer.
**SPEAKER_3** (4:35)
But any image stuff, any orchestration, distributed systems, GPUs?
**SPEAKER_4** (4:40)
No, I was doing basically nothing interesting.
**SPEAKER_3** (4:46)
Crud, web development?
**SPEAKER_4** (4:47)
Yeah, well, not web development. Just some basic, maybe some basic automation stuff.
**SPEAKER_3** (4:53)
Okay.
**SPEAKER_4** (4:54)
Just, yeah, no big companies or anything.
**SPEAKER_3** (4:59)
Yeah, but already some interest in automations, probably a lot of Python.
**SPEAKER_4** (5:03)
39 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000682691125