**Alessio** (0:05)
Hey, everyone, welcome to the Latent Space Podcast. This is Alessio, partner, and CTO and resident of Decibel Partners, and I'm joined by my co-host, Swix, founder of Small AI.
**Swyx** (0:15)
Hey, so today we're in the remote studio with Mikey Shulman, welcome.
**Mikey Shulman** (0:19)
Thank you, it's great to be here.
**Swyx** (0:21)
So I'd like to go over people's background on LinkedIn, and then maybe find out a little bit more outside of LinkedIn. You did your bachelor's in physics, and then a PhD in physics as well. Also, before going into Kensho Technologies, the home of a lot of top AI startups, it seems like, where you're head of machine learning for seven years.
You're also a lecturer at MIT, we talked about that, like what you talked about. And then about two years ago, you left to start Suno, which is recently burst on the scene as one of the top music generation startups. So we can go over that bio, but also I guess what's not on your LinkedIn that people should know about you?
**Mikey Shulman** (0:59)
I love music. I am an aspiring mediocre musician. I wish I were better, but that doesn't make me not enjoy playing real music.
And I also love coffee. I'm probably way too much into coffee.
**Alessio** (1:12)
Are you one of those people that, you know, they do the TikToks, they use like 50 tools to like grind the beans and then like brush them and then like spray them. Like whatever are we talking about here?
**Mikey Shulman** (1:22)
I confess there's a spray bottle for beans in the next room.
There is one of those weird comb tools. So guilty. I don't put it on TikTok though.
**Alessio** (1:33)
Yeah, no, no, some things got to stay. Got to stay private.
What do you play?
**Mikey Shulman** (1:37)
I played a lot of piano growing up and I play bass and I, in a very mediocre way, play guitar and drums.
**Alessio** (1:43)
That's a lot. I cannot do any of those things.
As Sean mentioned, you guys kind of burst into the scene as maybe the state of the art music generation company. I think it's a model that we haven't really covered in the past. So I would love to maybe for you to just give a brief intro of how do you do music generation and why is it possible? Because I think people understand you take texts and you have to predict the next word.
And you take a diffusion model and you basically add noise to an image and then kind of remove the noise. But I think for music, it's hard for people to have a mental model. How do you turn a music model on? What does a music model do to generate a song? So maybe we can start there.
**Mikey Shulman** (2:23)
Yeah, maybe I'll even take one more step back and say, it's not even entirely worked out.
I think the same way it is in text, and so it's an evolving field. If you take a giant step back, I think audio has been lagging images in text for a while. So I think very roughly you can think audio is like one to two years behind images in text. And so you kind of have to think today, like text was in 2022 or something like this. And the transformer was invented, it looks like it works, but it's far, far less established. And so, you know, I'll give you the way we think about the world now, but just with a big caveat that I'm probably wrong if we look back in a couple of years from now. And I think the biggest thing is you see both transformer based and diffusion based models for audio. And in ways that that is not true in text. I know people will do some diffusion for text, but I think nobody's like really doing that for real. So we prefer transformers for a variety of reasons. And so you can think it's very similar to text. You have some abstract notion of a token and you train a model to predict the probability over all of the next token. So it's a language model.
You can think anything language model is just something that assigns likelihoods to sequences of tokens. Sometimes those tokens correspond to text. In our case, they correspond to music or audio in general.
And I think we've learned a lot from our friends in the text domain from the pioneers doing this of how well these transformer models work, where do they work, where do they not work. But at its core, the way we like to do things with transformers is exactly like it works in text. Let me predict the next tiny little bit of audio, and I can just keep doing that and doing that and generating audio as long as I want.
41 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000649219483