**swyx** (0:03)
All right, we're here in the studio with Eiso Kant from Poolside together with Vibhu.
**Eiso Kant** (0:07)
Welcome. Thanks. Thanks for having me, guys. Good to be.
**swyx** (0:10)
Yeah, fresh on the plane. You texted me, you're like, hey, I'm on my way to SF. I was like, you're on the plane right now, right?
**Eiso Kant** (0:16)
Like, hey, you know, after I texted you, I realized that probably coming in with major jet lag is going to offer some fun experiences today, but let's do it.
**swyx** (0:24)
I mean, I think the thing I would tell the guests is that they don't actually have to prepare that much because if you're truly working on this every single day, then even like what you hazily remember is going to be new for a lot of the audience that don't live in your world every day, right? So 10 years ago, you did a talk at Google Slush talking about the democratization of AI.
And now here you are, like open sourcing an incredible new model that we're going to talk about. But I guess, what got you into democratization of AI? Like it's not obvious from your LinkedIn or something.
**Eiso Kant** (0:58)
No, it's not at all. Actually, I don't think it's obvious how I got in this space.
I owe getting into this space to Andrej Karpathy. In 2015, he wrote an article called The Unreasonable Effectiveness of Recurrent Neural Nets. And that article, I read it, and I pivoted my startup at the time overnight to working on RNNs and later LSTMs and Transformer models to be able to write code. If you go to this article and you scroll down, you can kind of start seeing like this was the precursor to what ended up becoming language models. So at least when he was character level language models that were starting to actually predict letters, he has an example out here. There's a little Paul Graham generator. And you can kind of read it and the text kind of makes sense, but it doesn't. And there's a little, there's an example of code a little bit further down.
**Vibhu** (1:45)
It's a Shakespeare.
**Eiso Kant** (1:49)
And for some reason, I read this and I went down the rabbit hole of learning everything I could about RNNs and LSTMs, right? This is pre-transformer paper. And I had built a completely unreasonable belief that neural nets should be able to generalize to anything and everything. And that language should be able to generalize to a lot of things that are intelligence and the ability to write code. And so I started building sourced, which was a fully open-source company, trying to build what we used to call machine learning on code, language models on code. And we spent about four or five years on this till the end of 2019 And that sounds really cool today, but back then, no one cared, right? Like no one cared. We were in the dark. Like we did things along the way. We tried applying convolutional neural nets to like the structure of code. We were, you know, when attention came out, we were applying it to LSTMs, and then the Transformer paper came out. And it was, it wasn't obvious. And what we missed throughout that entire journey, that we were on the right track, but we should have just kept scaling up. And today to all of us, the scaling laws and scaling up seems like the most obvious thing. But having spent four or five years of my life on working on language models on code, it wasn't obvious. I have a lot of respect to folks at Google and OpenAI and others who kind of took that confidence and kept going. We filled ultimately at the time and it kind of was like the biggest failure of my career, right? You blew $12 million of investors' money, which was a lot back then. Yep. You spent still a lot, but and you spent years with like a group of 40 people just obsessing over this problem. And life took a different turn and family kind of became the focus. And I kind of kept my head down. I'm running frankly didn't really look at language models for the following two years. Kind of big mistake considering. Following years is going to be really interesting. And then ChatGPT came out.
And it was kind of like a vindication. It's like people started texting me. I found like my old work decks and these old talks.
And throughout that whole journey, you know, we kind of really had a strong point of view at the time that like, as you're building more capable intelligence, it should be open and open source. When we started Poolside, that actually wasn't the case at all. And I want to be very open about it. When we started Poolside, we were like, there was a premise of two things. One is this technology is not going to stop compounding in capabilities. I think to most people obvious today, but three plus years ago when we started, most people were still arguing if these were stochastic parrots or not.
106 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000777982857