#49 - Microbes, Robots, and Ambition - Robin Sloan on His Novel Sourdough artwork

#49 - Microbes, Robots, and Ambition - Robin Sloan on His Novel Sourdough

Y Combinator Startup Podcast

November 29, 2017

Robin Sloan is a writer and media inventor based in Oakland.He just released his second novel, Sourdough.Kat Manalac is a Partner at YC.The YC podcast is hosted by Craig Cannon.
Speakers: Craig Cannon, Robin Sloan, Kat Manalac
**Craig Cannon** (0:00)
Hey, this is Craig Cannon, and you're listening to Y Combinator's podcast. Today's episode is with Robin Sloan and Kat Manalac. Robin's a writer and media inventor based in Oakland, and Kat's a partner here at YC.
So this October, Robin released his second novel, Sourdough, and a couple years ago, he released his first novel, Mr. Penumbra's 24-hour Bookstore.
So before we get going, if you could take a minute to review the podcast, that'd be awesome. All right, here we go. So this is kind of a weird jumping off point, but I listened to you on, I think it was a Mother Jones podcast, and you very briefly mentioned a machine learning experiment for the audio book. Could you talk about that a little bit longer?

**Robin Sloan** (0:38)
Sure, yeah, of course.
Well, okay, so the background is that as I've been working on these books that are in a lot of ways traditionally published, even though I have an interesting, very sort of forward thinking publisher, MCD, they still get printed on paper and sold mostly in bookstores and online and places like that.
As I've been working on all that stuff for a few years now, I've also, like many people in this area, many people that I'm sure you guys know, I've been really interested to the point of sort of preoccupation with machine learning, in particular the creative applications, like less the super practical sort of uses and the ways that it might transform the economy and all that. And I mean, truly more like the ways that we can use some of these systems to mash and mangle and just interact with like words and pictures and sounds in different ways. So that's all preface.
The audio book. First of all, it's actually interesting to know for folks that aren't totally plugged into the publishing industry, audio books are huge. It's like they're growing like gangbusters. Every time I go to a bookstore and do a reading, I ask people like, how many of you also listen to audio books? And I mean, truly everybody's hand goes out. It's like just a really, really popular way to consume this media. So kind of in step with that, audio book producers have gotten really serious and frankly, a little bit demanding. They're like, okay, Mr. Sloan, it's time to produce the audio book. What do you got for me? You know, like what can you produce or what will you produce that will make this a little distinct from just the printed book? They don't wanna just do like a recitation of what's on the page. So for this book, Sourdough, the story happens to hinge on this sourdough starter, you know, this little funny community of microbes that you use to bake this delicious bread. And in the story, there's a starter with some strange properties. And there's also this singing. There's like this music that, I don't wanna give any spoilers, but it kind of helps the starter grow, and it's all part of this mysterious package. So I described the music over and over in the book, I mean, like at great length, like, oh, it's so slow and sad and mysterious, and it's like in a language that no one understands. And so come time to produce the audio book, I was like, if you go through this whole thing in your earbuds, and you never hear even a scrap of this music, like, it's gonna be kind of disappointing.
Unfortunately, we did not have the budget to like hire the people who invent Dothraki for Game of Thrones or Klingon or whatever.
So we need some other way to like synthesize this sound that would be truly alien, like it's not any language on planet earth, it's something fictional, something invented. So this is where it loops back around to that obsession with machine learning, as I think you guys probably know, one of the things that these models can do really well is sort of take a corpus of stuff of training material and extract some patterns, some more general patterns, and then use those to generate something new but different, not just kind of mimicry of what you put in.
So I mean, at that point, it actually took a lot of kind of learning and tinkering with the code, and actually struggling with the code to get to this point. This was in, actually, funnily enough, this program was in Lua, in Torch, the original Torch, and totally a testament to just the power of this open source ecosystem. I mean, this was a paper written by one group of researchers implemented by this rogue, mad, machine learning genius in the UK, this guy named Richard Assar, who's just like, I bow down to him and his generosity truly in making this really wonderful and very usable implementation of this tool. It's called Sample RNN. And it takes as many MP3s as you want to feed it, chops them up into bits, churns learning for days and days and days, at least on my deep learning rig. I'm sure Google would be like, got it. For me, it took a few days. And then in the end, spits out this really, to my ear, at least weird and lovely kind of generalization. It tries its hardest to learn the essence of that music you fed it. Of course, it kind of fails, because none of these models are actually that good yet. At least it's stuff of that level of complexity. But the way in which it fails is really interesting. So that's all to say that now in this audio book, there is just these little whispers of this fictional music in this fictional language. And to my knowledge, it's the first time that the creative output of a machine learning system has been included in an audio book.

53 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000395407243