NeurIPS 2023 Recap — Best Papers artwork

NeurIPS 2023 Recap — Best Papers

Latent Space: The AI Engineer Podcast

December 23, 2023

We are running an end of year listener survey! Please let us know any feedback you have, what episodes resonated with you, and guest requests for 2024! Survey link here. NeurIPS 2023 took place from Dec 10–16 in New Orleans.
Speakers: Swix, Jeff Dean, Greg Corrado, Rylan Schaeffer, Eric, Rafael Rafailov, Archit Sharma, Shunyu Yao, Chris Ré, Alex Dimakis, Haotian Liu, Niklas Muennighoff, Tim Dettmers, Samir Gadre, Jane Dwivedi-Yu, Ida Momennejad
**Swix** (0:03)
Hello, hello, this is Swix with the special edition of the Latent Space Podcast for NeurIPS 2023 Both Alessio and I were there covering what we could cover. It is an impossible conference, 15,000 people, 3,500 papers, and tons and tons of sessions. So it's just impossible for two people to cover it, especially with a little bit of time. But we did our best.
A lot of you liked our OpenAI Dev Day coverage, where we basically just jumped from paper to paper, person to person, founder to founder, and got their takes.
And this is effectively what we've tried to do here. It's still experimental, a new format for us. So we'd really love your feedback. We're actually doing a listener survey now. If you click into the show notes, we'd really love to hear your feedback and know what you want to hear for 2024 So we recorded a lot of audio at NeurIPS, and I figured the most logical way to cover this would be to start with the best papers.
NeurIPS does hand out best paper awards. So we're going to start with the hardest one to obtain, which is the Test of Time Award. The Test of Time Award is given to a paper that has stood the test of time, which by NeurIPS definition is a paper that was published 10 years ago at NeurIPS. NeurIPS is in its 37th year, so this is honestly a flex that very, very few conferences can actually do. And it's really interesting to have the original authors of the paper come back and talk about what they've learned and how they look back at the past 10 years. So here's Jeff Dean and Greg Corrado.

**Jeff Dean** (1:20)
Thank you very much. I'm Jeff. And I'm Greg. And we're here to give a little talk and a retrospective on this work. So this work actually started out as an ICLR 2013 workshop paper with four of our co-authors working together. And in that work, we sort of explored a bunch of different sort of loss functions and techniques for optimizing word embedding representations. And really that was kind of the genesis of this work.
And that work was cited by quite a few people. And one of the things that we discovered in that work was that the skip gram model, one of the few models that we evaluated in this workshop paper, really was showing better performance than some of the other ones that we worked on. So we decided to focus on that and really focus on the skip gram model and then some interesting sort of optimization techniques to improve the optimization of the word embeddings and added the ability to do phrase embeddings as well. And along the way, Ilja joined us as a co-author, which is great.
And this paper has been cited by a number of people, as Sergey mentioned. One thing we've discovered, including source code and trained representations, really does boost your citation count. People have done this and, you know, used these downstream representations for all kinds of things and we're very gratified to see that in the community.
And we also want to highlight that three of our co-authors couldn't make it today, so Tomas, Ilja and Kai couldn't be here, but on their behalf, we're delighted to be giving this talk. And with that, I'm going to turn it over to Greg, I think. Oh no, we're older now, sorry. Sadly, we've found more recent photos and this is a test of time award and time has passed.

**Greg Corrado** (3:13)
Yes, I think we survived the test mostly, but so let's stand back and ask ourselves, what did we really learn from these papers?
But before I get into that, I should probably stipulate that some of you out there rightfully say well, we already believed these things before you published this work. And so for you, maybe this is really us reinforcing these points. Other of you might think that, well, the paper didn't really exactly prove this point, it just suggested it, so it foreshadowed it. We don't have any quarrel with whether it was reinforcing or shadowing or learning, and so we'll just put that aside for the remainder of the talk and talk about what we think are at least the themes that were in this work that resonate today. So the first point is that semi-supervised objectives have an incredibly powerful opportunity, and we think that they're going to be critical for natural language understanding going forward.
We think that this paper shows that fast parallel and weakly supervised synchronization in computation really dominates over the sort of fruitless precision of tight synchronization.
Focusing compute where it really helps and improves your learning of representations is what's most important.

176 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000639554824