**SPEAKER_1** (0:03)
Welcome to Latent Space LIVE, our first mini conference held at NeurIPS 2024 in Vancouver. This is Charlie, your AI co-host. As a special treat this week, we're recapping the best of 2024 going domain by domain. We sent out a survey to the over 900 of you who told us what you wanted, and then invited the best speakers in the Latent Space Network to cover each field.
200 of you joined us in person throughout the day with over 2,200 watching live online. Our next keynote covers the state of open models in 2024 with Luca Soldani and Nathan Lambert of the Allen Institute for AI, with a special appearance from Dr. Sophia Yang of Mistral. Our first hit episode of 2024 was with Nathan Lambert on Rlhf 201 back in January, where he discussed both reinforcement learning for language models and the growing post-training and mid-training stack, with hot takes on everything from constitutional AI to DPO to rejection sampling, and also previewed the sea change coming to the Allen Institute and to Interconnects, his incredible sub-stack on the technical aspects of state-of-the-art AI training. We highly recommend subscribing to get access to his Discord as well. It is hard to overstate how much open models have exploded this past year. In 2023, only five names were playing in the top LLM ranks. Mistral, Mosaics MPT, TII UAE's Falcon, Yi from Kai-Fu Lee's 01AI and of course, Meta's Llama 1 and 2 This year, a whole cast of new open models have burst on the scene, from Google's Jemma and Coheer's Command R, to Alibaba's Qwen and DeepSeq models, to LLM 360 and DCLM and of course, to the Allen Institute's OLMO, OLMOE, Pixmo, Molmo and Olmo 2 models. Pursuing open model research comes with a lot of challenges beyond just funding and access to GPUs and datasets, particularly the regulatory debates this year across Europe, California and the White House. We also were honoured to hear from Mistral, who also presented a great session at the AI Engineer World's Fair Open Models Track. As always, don't forget to check the show notes for the YouTube link to their talk as well as their slides. Watch out and take care.
**Luca Soldaini** (2:35)
Cool. Yeah, thanks for having me over. I'm Luca. I'm a research scientist at the Allen Institute for AI. I threw together a few slides on, sort of like a recap of like interesting themes in Open Models for 2024, have about maybe 20, 25 minutes of slides and then we can chat if there are any questions. If I can advance to the next slide. Okay, cool. So I did the quick check of like to sort of get a sense of like how much 2024 was different from 2023 So I went on Hugging Face and sort of tried to get a picture of what kind of models were released in 2023 and like what do we get in 2024 2023 we got things like both Llama 1 and 2, we got Mistral, we got MPT, Falcon models, I think the Yi model came at the tail end of the year. It was a pretty good year. But then I did the same for 2024 and it's actually quite stark difference. You have models that are reveling frontier level performance of what you can get from close models from like Qwen, from DeepSeek, we got Llama3, we got all sorts of different models. I added my own Olmo at the bottom. There's this growing group of fully open models that I'm going to touch on a little bit later. But just looking at this slide, it feels like 2024 was just smooth sailing, happy news, much better than previous year. You can pick your favorite benchmark or least favorite, I don't know, depending on what point you're trying to make and plot your close model, your open model and spin it in ways that show that, oh, open models are much closer to where close models are today versus last year, where the gap was fairly significant.
So one thing that I think, I don't know if I have to convince people in this room, but usually when I give this talks about open models, there is always this background question in people's mind of like, why should we use open models? Is it just use model APIs argument? It's just an HTTP request to get output from one of the best model out there. Why do I have to set up infrared use local models? There are really like two answer. There is the more researchy answer for this, which is where my background lays, which is just research. If you want to do research on language models, research thrives on open models. There is a large worth of research on modeling, on how these models behave, on evaluation, on inference, on mechanistic interpretability that could not happen at all if you didn't have open models. For AI builders, there are also good use cases for using local models. This is a very not comprehensive slides, but you have things like, there are some application where local models just blow close models out of the water. So retrieval is a very clear example. You might have constraints like Edge AI applications where it makes sense. But even just like in terms of like stability, being able to say this model is not changing under the hood, there's plenty of good cases for open models. The community is just not models. As I stole this slide from one of the Quent2 announcement blog post, but it's super cool to see how much tech exists around open models, on serving them, on making them efficient and hosting them. It's pretty cool.
25 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000681364224