2024 in AI Startups [LS Live @ NeurIPS] artwork

2024 in AI Startups [LS Live @ NeurIPS]

Latent Space: The AI Engineer Podcast

December 21, 2024

Happy holidays! We’ll be sharing snippets from Latent Space LIVE! through the break bringing you the best of 2024 from friends of the pod!
Speakers: Sarah Guo, Pranav Reddy
**SPEAKER_1** (0:15)
Welcome to Latent Space LIVE, our first mini conference held at NeurIPS 2024 in Vancouver. This is Charlie, your AI co-host. When we were thinking of ways to add value to our academic conference coverage, we realized that there was a lack of good talks just recapping the best of 2024 going domain by domain. We sent out a survey to the over 900 of you who told us what you wanted, and then invited the best speakers in the Latent Space Network to cover each field. 200 of you joined us in person throughout the day, with over 2,200 watching live online. Today, we're kicking it off with a keynote on the State of AI Startups. Sarah Guo, founder at Conviction and host of No Priors Podcast, and Pranav Reddy, partner at Conviction and former engineer at Neva. They'll be unpacking the top five themes of 2024 What ideas are working and what's not? From shifting market opportunities to why the supposed advantages of big tech incumbents might not be as strong as they seem. As always, don't forget to check the show notes for the YouTube link to their talk as well as their slides. Watch out and take care.

**Sarah Guo** (1:30)
Hi, everyone. My name is Sarah Guo. And thanks to Sean and friends here for having me and Pranav. So I start by just giving 30 seconds of intro. I promise this isn't an ad. We started a venture fund called Conviction about two years ago. Here is a set of the investments we've made. They range from companies at the infrastructure level in terms of feeding the revolution to foundation model companies, alternative architectures, domain-specific training efforts, and of course, applications. And the premise of the fund, Sean mentioned I worked at Greylock for about a decade before that and came from the product engineering side.
Was that we thought that there was a really interesting technical revolution happening. That it would probably be the biggest change in how people use technology in our lifetimes. And that represented huge economic opportunity. And maybe that there would be an advantage versus the incumbent venture firms. In that when the floor is lava, the dynamics of the markets change, the types of products and founders that you back change. It's a lot for existing firms to ingest and a lot of their mental models may not apply in the same way. And so there was an opportunity for first principles thinking. And if we were right, we would do really well and get to work with amazing people. And so we are two years into that journey and we can share some of the opinions and predictions we have with all of you. Pranav is going to start us off.

**Pranav Reddy** (2:55)
So quick agenda for today. We'll cover some of the model landscapes and themes that we've seen in 2024 What we think is happening in AI startups and then some of our latent priors on what we think is working and investing. So I thought it would be useful to start from what was happening at NeurIPS last year in December 2023 So in October 2023, OpenAI had just launched the ability to upload images to ChatGPT, which means up until that moment, it's hard to believe that roughly a year ago, you could only input text and get text out of ChatGPT. The Mistral folks had just launched the MixedRoll model right before the beginning of NeurIPS. Google had just announced Gemini. I very genuinely forgot about the existence of Bard before making these slides. Europe had just announced that they were doing their first round of AI regulation but not to be their last. When we were thinking about what's changed in 2024, there's at least five themes that we could come up with that feel like they were descriptive of what 2024 has meant for AI and for startups. So we'd start with, first, it's a much closer race on the foundation model side than it was in 2023 So this is Elm Arena. They're asking users to rate the evaluations from generations from specific prompts. So you get two responses from two language models to answer which one of them is better. The way to interpret this is roughly 100 ELO difference means that you're preferred two-thirds of the time. And a year ago, every OpenAI model was more than 100 points better than anything else. And the view from the ground was roughly, OpenAI is the IBM. There is no point in competing. Everyone should just give up, go work at OpenAI, or attempt to use OpenAI models. And I think the story today is not that. I think it would have been unbelievable a year ago if you told people that A, the best model today on this, at least on this e-val is not OpenAI, and B, that it was Google would have been pretty unimaginable to the majority of researchers. But actually, there are a variety of proprietary language model options and some set of open source options that are increasingly competitive. And this seems true not just on the e-val side, but also in actual spend. And so this is ramp data. There's a bunch of colors, but it's actually just OpenAI and Anthropics spend. And the OpenAI spend at the beginning, at the end of last year, in November of 23, was close to 90% of total volume. And today, less than a year later, it's closer to 60% of total volume, which I think is indicative both that language models are pretty easy, APIs to switch out, and people are trialing a variety of different options to figure out what works best for them. Related, second trend that we've noticed is that open source is increasingly competitive. So this is from the Scale Leaderboards, which is a set of independent evals that are not contaminated. And on a number of topics that actually the foundation models clearly care a great deal about, open source models are pretty good on math, instruction following, and adversarial robustness. The Llama model is amongst the top three of evaluated models. I included the agenting tool used here just to point out that this isn't true across the board. There are clearly some areas where foundation model companies have had more data or more expertise in training against these use cases, but models are surprisingly and increasing, open source models are surprisingly increasingly effective. This feels true across evals. This is the MMLU eval. I want to call it two things here. One is that it's pretty remarkable that the ninth best model and two points behind the best state-of-the-art model is actually a 70 billion parameter model. I think this would have been surprising to a bunch of people where the belief was largely that most intelligence is just an immersion property and there's a limit to how much intelligence you can push into smaller form factors. In fact, a year ago, the best small model or under 10 billion parameter model would have been Mistral 7B, which on this eval, memory serves somewhere around a 60 Today, that's the Llama 8B model, which is more than 10 points better. The gap between what is state-of-the-art and what you can fit into a fairly small form factor is actually shrinking.

40 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000681199633