ICLR 2024 — Best Papers & Talks (ImageGen, Vision, Transformers, State Space Models) ft. Durk Kingma, Christian Szegedy, Ilya Sutskever artwork

ICLR 2024 — Best Papers & Talks (ImageGen, Vision, Transformers, State Space Models) ft. Durk Kingma, Christian Szegedy, Ilya Sutskever

Latent Space: The AI Engineer Podcast

May 27, 2024

Speakers for AI Engineer World’s Fair have been announced! See our Microsoft episode for more info and buy now with code LATENTSPACE — we’ve been studying the best ML research conferences so we can make the best AI industry conf!
Speakers: Charlie, Timothée Darcet, Durk Kingma, Pablo Pernías, Hila Chefer, Ilya Sutskever, Christian Szegedy, Sachin Goyal, Pulkit Tandon, Shashank Venkataramanan, Yukang Chen, Bowen Peng, Suyu Huang, Guanhua Wang, Sasha Rush
**Charlie** (0:12)
Welcome to the Latent Space Podcast, ICLR edition. This is Charlie, your AI co-host. This month, we attended the 12th International Conference on Learning Representations in Vienna, Austria. ICLR is a newer conference, but already gaining a lot of popularity as a deep learning focused academic research conference, roughly half the size of Neurope's. Many of you absolutely loved our NeurIPS coverage last year. And while we can't do that for every conference, we're proud to bring you a special two-part episode covering our attempt at giving you an audio experience of ICLR.
If you'd like to see us return to Vienna for ICML, let us know by sharing this episode on X and LinkedIn. For our subset of AI engineering concerns, which you can see on the LatinSpaceAbout page, we're saving agents and reasoning topics and introducing everything else we saw at ICLR.
This episode covers the best papers across Section A, Image Generation and Diffusion, Section B, Computer Vision and Weak Supervision, Section C, Improving Attention Algorithms, Section D, State Space Models and the rest of the poster sessions we saw. This is the first of two episodes covering ICLR and will be overwhelmingly academia-focused.
If you're interested in production AI engineering and industry, you should join us at the first AI Engineer World's Fair this June, where we have now announced many of our speakers from all the big clouds including Microsoft Azure AI and GitHub CEO Thomas Domka. All the large model labs including OpenAI, DeepMind, Mistral and Adept. All top AI-enabled developer tools and Codigen agents including fan favorite guest Chris Lattner of Modular and Scott Wu of Cognition Labs Devon. Major GPU and inference providers like NVIDIA, Grok, Fireworks and upcoming guest gradient AI. All the rest of the emerging LLM OS stack of startups and open source tools across Reg, Multi Modality, LLM Ops and agent frameworks like Instructor, LangChain, LaMendex, DSPy, Unsloth, Crew AI, disruptive startups like Mid Journey, Perplexity and Character AI and for the first time, talks about AI deployed at massive scale from Salesforce to Novartis, to Tinder, to Coinbase, to Khan Academy. Get your tickets now and see you in San Francisco from June 25th to 27th. We'll start this episode with an extended meditation on what deep learning representations really entail. The inaugural ICLR test of time award went to Kingma and Welling for auto-encoding variational bays. The paper that introduced the variational autoencoder, a key precursor to diffusion models.
But first, we'll introduce what VIEs are and then have Durk Kingma talk about his 10-year retrospective on the VIE in his test of time award speech. But first, the best way to introduce VIEs is to start with a clip from the Archive Insights YouTube channel, which we've linked to in the show notes. We start from a basic knowledge of basic autoencoders, build up to denoising autoencoders, and then the key things to know about variational autoencoders. Watch out and take care.

**SPEAKER_2** (3:40)
Okay, so that's the basic idea behind autoencoders, but there are a few very clever tricks that you can apply to an autoencoder to have it do some really fancy stuff. So imagine that you start with a normal MNIST digit. It's a clean image, nothing's wrong with it. But then you add a whole lot of noise to it, and you're going to run that noisy image through your encoder network.
You get through the bottleneck representation, and then you try to reconstruct the image, but instead of reconstructing the noisy image, what you're going to do is try and reconstruct the original clean image.
And if you train this network in a whole bunch of these noisy MNIST digits, you're going to try and force the encoder step to actually get rid of the noise. And this is what we call a denoising autoencoder. And so you can see here that by using this approach, you can actually train a denoising autoencoder that is very good at removing noise from input images. And denoising images isn't the only thing that you can do with this type of approach. So in this case, for example, you take an input, and instead of adding noise to it, you simply crop a rectangular area out of the image and you throw it away.

**SPEAKER_3** (4:38)
You replace it with white or black pixels.

**SPEAKER_2** (4:41)
You feed that input image through the network, and you try to reconstruct the original full image. And this technique is what we call neural inpainting. It's where you take a small part of the image, you throw it away, and then you ask the network to reconstruct whatever was there in the input image. And with this approach, you can do simple things like removing watermarks from images, but you could also remove a parked car, for example, if you are filming on a movie set in a natural setting.

175 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000656912933