**SPEAKER_2** (0:12)
So welcome to Latent Space. We are basically trying to provide the best optimal sort of podcast experience of NeurIPS for people who are not here. And congrats on your paper. How does it feel?
**Kevin Wang** (0:22)
Yeah, it was very exciting. We had a poster yesterday and then today we'll have an oral talk.
**SPEAKER_2** (0:27)
Were you just like mobbed?
**Kevin Wang** (0:29)
There was a lot of people. It's like three hours straight of like waves of people to like that we were trying to stupid.
**SPEAKER_2** (0:35)
So I've never received the best paper. Did you just find out on the website? Like what?
**Kevin Wang** (0:40)
I just like woke up one day and like checked my email. And then they just told me, yeah, they was like, oh, like that's like, I saw email. You were like been awarded Best Paper.
**SPEAKER_2** (0:51)
Maybe you know from the reviews as well, right?
**Kevin Wang** (0:54)
We know from the reviews that we did well. But there's a difference between like doing well on the reviews and getting Best Paper. So right now we didn't actually know.
**SPEAKER_2** (1:02)
Yeah. Okay. So I skipped a little bit. Maybe we can go sort of one by one and sort of introduce, you know, who you are and what you did on the team.
**Kevin Wang** (1:11)
I'm Kevin. I was an undergrad from Princeton. I just graduated. And yeah, I guess I led the project, like started the project and then was very happy to collaborate with Ishaan and Nicole and Ben also.
**SPEAKER_2** (1:24)
Right. And were you in like the same research group?
**Kevin Wang** (1:26)
Like how do you, how does your social context? So yeah, so we're all from Princeton. Yeah.
**SPEAKER_2** (1:33)
Thanks to Allen for booking you guys.
**Kevin Wang** (1:35)
So this project actually started from like an IW seminar. So like independent work research seminar that Ben was teaching. And this was like actually like one of my first experiences in like ML research. So it was really valuable to get that experience. And then Ishaan was also in that seminar and working on adjacent things. So we collaborated a lot during that seminar. And then yeah, the project turned out to have some pretty cool results. And then later on also like the HALT working on sort of similar things also joined it on the project and became like a good collaboration.
**SPEAKER_2** (2:05)
Yeah. And I don't know if any of you guys want to chime in on like other elements of coming into like deciding on this problem.
**Benjamin Eysenbach** (2:14)
So it's like probably my lab works on deep reinforcement learning, but historically deep meant like two or three or four layers.
**SPEAKER_2** (2:23)
Not 1000
**Benjamin Eysenbach** (2:25)
When Kevin and Ishaan making the one to try really deep networks, it's kind of skeptical it was going to work. I've tried this before, it doesn't work. Other papers have tried this before and doesn't work. So I was very, very skeptical starting out. I don't know if I conveyed this at the time, but that was my prior going in because...
**SPEAKER_2** (2:41)
But do you view your job as like screening or like, hey guys, this is probably isn't going to work, you should try a different idea, you know, like, or should you be encouraging even if it's dumb?
**Benjamin Eysenbach** (2:50)
It's selecting bets. Yeah. And this was a bet I was willing to make.
**SPEAKER_2** (2:55)
What made you willing to make a bet?
**Benjamin Eysenbach** (2:57)
It seemed relatively low cost in that we, Michał, in particular, had spent the past year developing infrastructure that made it a lot easier to run some of these experiments. And the precedent was deeper networks should do a whole lot better. Like, that's what the deep learning revolution has been over the last day.
**SPEAKER_2** (3:15)
Yeah, I know. Why do we stop making them deeper?
**Benjamin Eysenbach** (3:17)
And reinforcement learning was like this one anomaly where we continue to use these really shallow networks. And that's particularly true in the settings that we were looking at, where you're starting from scratch, you're starting from nothing.
**SPEAKER_2** (3:27)
Any other perspectives you guys want to chime in with?
**Kevin Wang** (3:30)
I guess maybe I should just go over like an overview of our project.
**SPEAKER_2** (3:33)
Yes, okay, sorry.
**Kevin Wang** (3:35)
So the way that I kind of view our project is that if you look at the landscape of deep learning, you have NLP, like language, vision, and then RL. And as Ben kind of alluded to, in language, in vision, we've sort of converged to these paradigms of scaling to massive networks, right? Like hundreds of billions of parameters, trillions of parameters. And there's been a lot gained in deep learning from that, right? But then it seems like in the third sort of branch of deep learning in deep RL, that has not yet been the case. Like I was very surprised like coming into some like, you know, Ben's class and seminar, when I was looking at the networks, oh, why were you just using like a simple two layer MLP for like these frontier sort of, you know, state of the RL algorithms? And so I was very curious, like, can we design RL algorithms? Can we sort of put together a recipe for RL that can allow it to scale in potentially, you know, analogous ways that language and vision might scale? And so what we did is that we know that traditional RL, like, let's say, like, value-based RL doesn't really scale, right? This is pretty clear from the literature. So we tried a different approach to RL called self-supervised RL, where instead of learning, like, a value function, we're learning representations of states, actions, and future states, such that the representations along the same trajectory are pushed together, the representations along different trajectories are pushed apart. And this is just like a different approach to RL that allows us to learn in a self-supervised manner. So we can solve task reach goals without any human crafted reward signal. And so we know that self-supervised learning is scalable in these different areas in deep learning. So can self-supervised RL scale in similar ways? When we first tried it, it actually didn't work. We made the network steeper, the performance totally degraded. But then we also tried... But then I separately was like... There's also some other work, like in RL literature, we tried residual connections, and there's a few other architectural components that we had to put into the recipe. And then all of a sudden, one day, like I ran this experiment, and there was like this one environment in which there was like, like going from like, like doubling the depth didn't really do anything, but like doubling the depth again, with these different components, suddenly like skyrocketed performance in this one environment.
23 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000748427729