#45 - Building Dota Bots That Beat Pros - OpenAI's Greg Brockman, Szymon Sidor, and Sam Altman artwork

#45 - Building Dota Bots That Beat Pros - OpenAI's Greg Brockman, Szymon Sidor, and Sam Altman

Y Combinator Startup Podcast

November 8, 2017

Greg Brockman is the CTO and cofounder of OpenAI.Szymon Sidor is a Research Scientist at OpenAI.Sam Altman is the President of Y Combinator and Co-Chairman of OpenAI.Watch their bot compete at The International.
Speakers: Craig Cannon, Greg Brockman, Sam Altman, Szymon Sidor
**Craig Cannon** (0:00)
Hey, this is Craig Cannon, and you're listening to Y Combinator's podcast. Today's guests are Greg Brockman, Szymon Sidor and Sam Altman. Greg is the CTO and co-founder of OpenAI. Szymon is also at OpenAI, he's a research scientist there.
And before we get going, if you haven't yet subscribed or reviewed the podcast yet, that would be awesome if you did. All right, here we go.

**Greg Brockman** (0:21)
Now, if you look forward to what's gonna happen over upcoming years is the hardware for these applications for running neural nets really, really quickly, are gonna get fast, faster than people expect. And I think that what that's gonna unlock is you're going to be able to scale up these models and you're going to see qualitatively different behaviors from what you've seen so far. At OpenAI, we see this sometimes. For example, we had a paper on this unsupervised learning where you train a language model to, you train a model to predict the next character in Amazon reviews. And just by learning to predict the next character in Amazon reviews, somehow it learned a state-of-the-art sentiment analysis classifier.
And so it's kind of crazy if you think about it, right? You just were told, hey, predict the next character. You know, if you were told to do this, well, the first thing you do is you'd learn the spelling of words and you learn punctuation. The next thing you do is you start to learn semantics, right? If you have extra capacity there. And that this effect goes away if you use a slightly smaller model. And what happens if you have a slightly larger model? Well, we don't know because we can't run those models yet.
But in upcoming years, we'll be able to.

**Sam Altman** (1:25)
What do you guys think are the most promising and underexplored areas in AI? If we're trying to make it come faster, what should people be working on that they're not?

**Szymon Sidor** (1:34)
Yeah, so there are many areas of AI that we already developed by quite a bit. There's some basic research in just classification, deep learning and reinforcement learning. And what people do is people kind of try to invent problems and such as solving some complicated games of hierarchical structure, and they try to add kind of extra features to their models to combat those problems. But I think there's very little research happening on actually understanding the existing methods and their limits.
So, for example, it was a long-held belief in deep learning that to parallelize your computation, you need to cram as small batches as possible on every device. And in fact, Baidu did this impressive engineering feat where they took recurrent neural networks and they implemented the kind of GPU assembly code to make sure that you can fit batch size one RNNs on every GPU. And despite all those smart people working on this problem, only very recently did Facebook kind of took a cold, quiet look at just very basic problem of classification.
And in their great paper called Image Let It One Hour, they showed that if you actually take a code that does image classification, and if you fix all the bugs, you actually can get away with much larger batch size and therefore finish the classification problem much faster. And it's not the kind of sexy research that people want to see where you have some hierarchy of big, I don't know, but actually this kind of research, I think at this point will advance field the most.

**Craig Cannon** (3:31)
So Greg, you mentioned hardware in your initial answer. In the near term, what are the actual innovations that you foresee happening?

**Greg Brockman** (3:39)
So the big change is that the kinds of computers that we've been trying to make really fast are general purpose computers that are built on the von Neumann architecture. You basically have a processor, you have a big memory, and you have some bottleneck between the two. With the applications that we're starting to do now, suddenly you can start making use of massively parallel compute.
The architectures that these models can run on the fastest are going to look kind of like the brain, where the brain is basically you have a bunch of neurons that all have their own memory right near to them, and that they all talk to their neighbors, and maybe there's some kind of longer-range skip connections, and that just no one's really had incentive to develop hardware like this. And so what we've seen is that, well, you move your neural networks from running on a CPU to a GPU, and now suddenly you have a thousand CUDA cores running in parallel, and that you can get massive performance boost there. Now if you move to specialized hardware that is sort of much more brain-like and that runs a bunch of, you know, that sort of runs in parallel with a bunch of tiny little cores, that you're going to be able to run these models sort of insanely faster.

57 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000394558264