ICLR 2024 — Best Papers & Talks (Benchmarks, Reasoning & Agents) — ft. Graham Neubig, Aman Sanger, Moritz Hardt) artwork

ICLR 2024 — Best Papers & Talks (Benchmarks, Reasoning & Agents) — ft. Graham Neubig, Aman Sanger, Moritz Hardt)

Latent Space: The AI Engineer Podcast

June 10, 2024

Our second wave of speakers for AI Engineer World’s Fair were announced! The conference sold out of Platinum/Gold/Silver sponsors and Early Bird tickets! See our Microsoft episode for more info and buy now with code LATENTSPACE.
Speakers: Charlie, Graham Neubig, Hunter Lightman, Noam Brown, Moritz Hardt, Carlos Jimenez, Thomas Scialom, Yonatan Oren, Akari Asai
**Charlie** (0:05)
Welcome to the Latent Space Podcast, ICLR Edition, Part 2 This is Charlie, your AI co-host. We're back with our coverage of the 12th International Conference on Learning Representations in Vienna, Austria. Many of you absolutely loved our NeurIPS coverage last year, and we're proud to bring you Part 2 of our special two-part episode covering our attempt at giving you an audio experience of ICLR. If you'd like to see us return to Vienna for ICML, let us know by sharing this episode on X and LinkedIn. In Part 1, we covered the best papers of ICLR across four sections. Image generation and diffusion, computer vision and weak supervision, improving attention algorithms and state space models.
In today's episode, we cover the wealth of papers we found around the related problems of LLM reasoning and agents, also in four sections. In Section A, we do a regular latent space chat with Graham Neubig of Web Arena and Opendevin, introducing many of the major themes of the rest of this episode. In Section B, we survey a few prominent issues in benchmarking from SUEbench, test set contamination and general intelligence. In Section C, we look at agent building blocks from rag and self-reflection, verification, safety and frameworks. In Section D, we finally look at two proposed agent systems from Google DeepMind's web agent and MetagePete.
This is the second of two episodes covering ICLR, which is overwhelmingly academia focused. If you're interested in production AI engineering and industry, you should join us at the first AI engineer world's fair this June, where we have now announced many of our speakers from all the big clouds, all the large model labs, now including Anthropic, Cohere and Cartesia, the brand new state space model startup, all top AI enabled developer tools and code gen agents, now including Quinn Slack, CEO of Sourcegraph, insights on the GPU and inference market, like last week's guest gradient AI, and now featuring Dylan Patel of the Semi Analysis GPU Poor Blog, and the rest of the emerging LLM OS stack of startups and open source tools across rag, multimodality, LLM Ops and agent frameworks, disruptive startups like Mid Journey, perplexity and character AI, and for the first time, a closed door track for VPs of AI and technical leaders to discuss AI strategy and leadership. Get your tickets now and see you in San Francisco from June 25th to 27th. We'll start this episode with a special on-site interview we did at ICLR with Professor Graham Neubig of the Language Technologies Institute of Carnegie Mellon University. Graham has taught the CMU NLP course for the last seven years, but is also an active participant in the open-source AI software ecosystem, having been personally involved in Opendevin, which recently scored a notable 21% unassisted resolve rate on SWE Benchlight.
As an extra special treat, we are proud to invite Aman Sanger, co-founder of Cursor AI, back as our first ever guest co-host to add his personal takes on the state of code editing and agents. This is going to be a doozy of an episode, so we better get started. Watch out and take care.

**SPEAKER_3** (3:43)
So, welcome to the pod, Graham, and welcome to the pod, Aman, our first ever guest co-host.

**Graham Neubig** (3:50)
Yeah, thanks for having me.

**SPEAKER_3** (3:51)
Yeah, thanks for taking some time during ICLR. This is like very impromptu, but the two of you wanted to chat. I was like, let's just record a chat, and that can be fun. And also, one of my goals here at a conference is like this, is just cover posters for people who are at home, like not at a conference like this, just to get a sense of the mood. So, I'll cover a little bit of your background, and then you can sort of do filling the blanks. So, you're a professor at CFU, teach NLP. You also spent some years in Japan as a language teacher as well as a grad student.

**Graham Neubig** (4:19)
Yeah, it was a good experience. It was very impromptu, me going there, just because I wanted to learn a new language, but I ended up staying there for 11 years.
Yeah, amazing. Never know how life leads you, I guess.

**SPEAKER_3** (4:29)
Yeah, now you're on your own lab, and you teach the advanced NLP course. I'm sure that's been a wild ride over the past seven years.
Was it, it was 2016?

**Graham Neubig** (4:38)
Yeah, the last time I taught it before this semester was before ChatGPT. So I had to go in and rip out everything and put all the new stuff in. But yeah, it's a good opportunity to keep up with all the new stuff too, because I feel like I need to pressure myself into giving a good experience. So I need to cover all the areas.

238 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000658401616