#20 - Making Music and Art Through Machine Learning - Doug Eck of Magenta artwork

#20 - Making Music and Art Through Machine Learning - Doug Eck of Magenta

Y Combinator Startup Podcast

July 21, 2017

Doug Eck is a research scientist at Google and he’s working on Magenta, a project making music and art through machine learning. If you want to learn more you can check out Magenta.Tensorflow.org
Speakers: Craig Cannon, Doug Eck
**Craig Cannon** (0:00)
Hey, this is Craig Cannon, and you're listening to Y Combinator's podcast. Today's episode is with Doug Eck. Doug's a research scientist at Google, and he's working on Magenta, which is a project making music and art through machine learning. Their goal is to basically create open source tools and models that help creative people be even more creative. So if you want to learn more about Magenta or get started using it, you can check out magenta.tensorflow.org. All right, here we go.
I wanted to start with the quote that you ended your IO talk with, because I feel like that might be helpful for some folks. So it's a Brian Eno quote, and I will have the slightly longer version. Yeah, good. So yeah, it goes like this.
Whatever you now find weird, ugly, uncomfortable, and nasty about a new medium will surely become its signature. CD distortion, the jitteriness of digital video, the crap sound of 8-bit, all of these will be cherished and emulated as soon as they can be avoided. It's the sound of failure. So much modern art is the sound of things going out of control. Out of a medium, pushing to its limits and breaking apart.
So that's how you ended your IO talk.

**Doug Eck** (1:02)
Correct.

**Craig Cannon** (1:03)
And what it kind of opened up for me was like what, when you're thinking about creating Magenta and all the projects therein as new mediums, how are you thinking about, like, how are you thinking about what's gonna be broken and what's gonna be created?

**Doug Eck** (1:20)
The reason that I put that quote there, I think, is to be honest with the division between engineering and research and artistry and to not think that what I'm doing is being a machine learning artist, but we're trying to build interesting ways to make new kinds of art.
And I think, you know, it occurred to me, I read that quote and I thought, you know, that's it, right? No matter how hard Eastman or whomever invented the film camera, I'm sorry if that's the wrong person, right, like, they clearly weren't thinking of breakage, or they're trying to avoid certain kinds of breakage.
I mean, you know, guitar amplifiers aren't supposed to distort, you know. And, you know, I thought, well, what if we do that with machine learning? Like, models that are, like, the first thing you're gonna do, if you think, if someone comes to you and says, here's this really smart model that you can make art with, what are you gonna do? You're gonna try to show the world that it's a stupid model, right? But maybe the way that, maybe it's smart enough that it's kind of hard to make it stupid, so you get to have a lot of fun making it stupid, right?

**Craig Cannon** (2:18)
I was playing with a quick draw this morning with my girlfriend, and what she was trying to do was make the most accurate picture that the computer wouldn't recognize, like immediately out of the gate. She works in art, and like, yeah, doesn't want to believe.

**Doug Eck** (2:31)
That's right. It's a good intuition, I mean, you know.

**Craig Cannon** (2:34)
Yeah, so maybe the best way to start is then talk about like, what are you working on right now? What are you guys making?

**Doug Eck** (2:41)
So right now, we're working on, kind of let me think, it's a good question.
We have this project called Ensynth, which is trying to get deep learning models to generate new sounds.
And we're working on a number of ways to make that better. I think one way to think about it is we have these, we have this latent space. We have a, so to make that a little bit less a buzzword, we have a kind of compressed space, a space that doesn't have the ability to memorize the original audio, but it's set up in such a way that we can try to regenerate some of that audio. And in regenerating it, we don't get back exactly what we started with, but hopefully we get something close. And that space is set up so that we can move around in that space and come to new points in that space and actually listen to what's there.
Right now, it's quite slow to listen, so to speak. We're not able to do things in real time. And we also would love to be at kind of a meta level, building models that can generate those embeddings, having trained on other data, so that you're able to move around in that space in different ways. And so we're kind of moving, we're continuing to work with sound generation for music. And we also are spending quite a bit of time on rethinking the music sequence generation work that we're doing. We put out some models that were, by any reasonable account primitive, I mean, kind of very simple recurrent neural networks that generate MIDI from MIDI, and that maybe use attention, that maybe have a little bit smarter ways to sample when doing inference, when generating. And now we're actually taking seriously, wait a minute, what if we really look at large data sets of performed music? What if we actually start to care about expressive timing and dynamics, care deeply about polyphony, and really care about not putting out kind of what you would consider a simple reference model, but actually what we think is super good. And I think those are the things we're focusing on. I think we're trying to actually make things really pull up quality and make things that are better and more usable for people.

39 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000390144250