How To Build The AGI Future: Bob McGrew artwork

How To Build The AGI Future: Bob McGrew

Y Combinator Startup Podcast

January 31, 2025

According to OpenAI's former Chief Research Officer Bob McGrew, reasoning and test-time compute will unlock more reliable and capable AI agents— and a path to scale to AGI.
Speakers: Bob McGrew, Garry Tan
**Bob McGrew** (0:00)
If you ask people what AGI was, they would say, it's a model that you can actually interact with, it passes the Turing test, it can look at things, it can write code, it can even draw an image for you.

**Garry Tan** (0:09)
We're there.

**Bob McGrew** (0:10)
Yeah, and we've had this for years. And if you said, okay, well, what happens when you get all those capabilities? They say, well, everybody's out of a job and game over for humanity. And none of that is happening. I think in the big picture, we're reaching that bottleneck for pre-training and data. But now we have this new mechanism with reasoning and test time compute. What we're going to see out of reasoning is that it's really going to unlock the possibility of agents to do actions on your behalf, which has sort of always been possible, but it's just never been quite good enough. And you really need a lot of reliability. I think that is now in sight.

**Garry Tan** (0:45)
Hey, guys, we have a real treat today. Bob McGrew, formerly Chief Research Officer at OpenAI. You were a part of building a lot of the research team. What was that like early at OpenAI?

**Bob McGrew** (1:00)
The really interesting thing about OpenAI is that I did not originally intend to go to a research lab. When I left Palantir, I wanted to start a company. I had a thesis that robotics would be the first real business that was built out of deep learning. This was back in 2015 And I talked my way into a friends nonprofit.
I never had a badge, but I would go in, he'd open the door for me. And I learned deep learning by teaching a robot how to play checkers from vision. And in the process of doing this, I learned a lot about robotics. And I learned that robotics was definitely not the right startup to start in 2015 or 2016 I ended up going to OpenAI basically because it was a place full of very smart people. And it had big ambitions. It was a place where I could really learn. I had all this management experience from Palantir, but it was just a place for me to really become an expert in deep learning. And from there, figure out what it could actually be used and applied for.

**Garry Tan** (2:04)
What were some of the earliest things that you remember working on? And how did that play into what everyone knows OpenAI to be now?

**Bob McGrew** (2:10)
Yeah, when OpenAI started, the goal was always to build AGI. But the theory early on was that we would build AGI by doing a lot of research and writing a lot of papers, and we knew that this was a bad theory. I think for a lot of the early people who were startup people, Sam, Greg, myself, it felt painful and a little academic. But at the same time, it was what we could do at the time. And so some of the early projects, I worked on a robotics project where we took a robot hand, a humanoid robot hand, and we taught it to solve Rubik's Cube. The idea in doing that was that if we could make the environments complicated enough, the artificial intelligence would be able to generalize out of the narrow domain it was taught, and learn something more complicated, which was one of the ideas that later we see coming back with LLMs. The other really early big project was solving Dota 2 So there's a long history of solving games as a path towards building better AI, from Othello to Go. And after beating Go, the next hardest set of games are actually video games. They're not very classy, but they're a lot of fun. And I can assure you that mathematically they were harder. And so DeepMind went after StarCraft, OpenAI went after Dota 2 And there was real insight that was generated there, which was that it really strengthened our belief that scale was the path to improving artificial intelligence. That with Dota 2, the secret idea was that we could take huge amounts of experience, and feed it into a neural network, and that the neural network would actually learn and generalize from that. And later, we actually went back and applied this to the robot hand, and that became the key idea for the robot hand. And at the same time as these two big projects were going on, Alec Radford was experimenting with language. And the core idea behind GPD-1 is that if you have a transformer and you apply this super simple objective of guessing the next token, guessing the next word, that that would be enough signal that you could actually have something that would be able to generate coherent text. And in retrospect, it sounds sort of obvious, right? Like, you know, clearly this was going to work, but no one thought this would work at the time. Alec, you know, really had to persevere for years in order to make this work. And that became GPD-1. And then after GPD-1 seemed successful, we brought in the ideas from Dota and from the robot hand of training at, you know, larger and larger amounts of scale, and training on a really diverse set of data and looking for generalization. And together, that brings you to GPD-2 and GPD-3 and GPD-4.

26 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000687512732