**Alessio Fanelli** (0:10)
Hey, everyone. Welcome to the Latent Space Podcast. This is Alessio, Partner and CTO and Residence at Decibel Partners, and I'm joined by my co-host, Zwick, Writer and Editor of the Latent Space. Up until today, we never verified that we're actually humans to you guys, so we'd have one good thing to do today would be run ourselves through some AI benchmarks and see if we are humans indeed.
So since I got you here, Sean, I'll start with one of the classic benchmark questions, which is what movie does this emoji describe?
The emoji set is little kid, blue fish, yellow blue fish, orange puffer fish. What movie does that describe?
**Swyx** (0:51)
I think if you added an octopus, it would be slightly easier, but I prep this question so I know it's finding Nemo.
**Alessio Fanelli** (0:57)
You are so far human.
Second one of these emoji questions instead depicts a superhero man, a superwoman, three little kids, one of them, which is a toddler. So you got this one too?
**Swyx** (1:11)
Yeah, it's one of my favorite movies ever. It's the Incredibles.
Second one was kind of a let down, but the first is a classic.
Okay, I'm going to ramp it up a little bit. So let's ask something that involves a little bit of world knowledge. So when you drop a ball from rest, it accelerates downward at 9.8 meters per second. If you throw it downward instead, assuming no air resistance, so you're throwing it down instead of dropping it. It's acceleration immediately after leaving your hand is A, 9.8 meters per second, B, more than 9.8 meters per second, C, less than 9.8 meters per second, D, cannot say unless the speed of the throw is given.
**Alessio Fanelli** (1:46)
I would say B. You know, I started as a physics major and then a science, but I think I got enough from my first year that is B.
**Swyx** (1:54)
You have even proven that you're human because you got it wrong, whereas the AI got it right. Is 9.8 meters per second the gravitational constant because you are no longer accelerating after you leave the hand. The question is, if you throw it downward after leaving your hand, what is the speed?
It goes back to the gravitational constant, which is 9.8 meters per second. I thought you said you're a physics major.
**Alessio Fanelli** (2:18)
That's why I changed. I'm a human.
But you got them all right.
**Swyx** (2:22)
I can't ramp it up. So assuming the AI got all that right, you would think that the AI would get this one wrong because it's just predicting the next token, right?
**Alessio Fanelli** (2:31)
Right.
**Swyx** (2:32)
In the complex z-plane, the set of points satisfying the equation z squared equals modulus z squared is A, a pair of points, B, a circle, C, a half line, D, a line. The processing is going on in your head.
**Alessio Fanelli** (2:51)
You got minus 3, a line? This is hard.
**Swyx** (2:55)
Yes, that is a line. What's funny is that I think if an AI was doing this, it would take the same exact amount of time to answer this as it would every single other word because it's computationally the same to them.
So anyway, if you haven't caught on, today we're doing our first AI fundamentals episode with just the two of us, no guess, because we wanted to go deep on one topic and the topic is?
**Alessio Fanelli** (3:18)
AI benchmarks.
**Swyx** (3:19)
So why are we focusing on AI benchmarks?
**Alessio Fanelli** (3:22)
So GPT-4 just came out last week, and every time a new model comes out, all we hear about is it's so much better than the previous model on benchmark X and benchmark Y. It performs better on this, better on that, but most people don't actually know what actually goes on under these benchmarks.
So we thought it would be helpful for people to put these things in context. And also benchmarks evolved. Like the more the models improve, the harder the benchmarks get. Like I couldn't even get one of the questions right. So obviously they're working. And you'll see that from the 1990s, where some of the first ones came out to the day, the difficulty of them has really skyrocketed. So we want to give a brief history of that and leave you with a mental model on, okay, what does it really mean to do well at X benchmark versus Y benchmark?
So excited to dive in.
**Swyx** (4:15)
Yeah. I would also say when you ask people, what are the ingredients going into a large language model? They'll talk to you about the data. They'll talk to you about the neural nets. They'll talk to you about the amount of compute, how many GPUs are getting burned based on this.
43 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000607803441