**Unknown Host** (0:00)
Hi, listeners. As you may know, I recently wrapped up the AIE Code Conference in New York, and while I'm traveling, I do like to visit top AI startups in person to bring you interviews that you don't find on any other podcast that just does a Zoom call. General Intuition, or GI for short, is a spin-out of a 10-year-old game clipping company called METAL, which has 12 million users, but in comparison, Twitch only has 7 million monthly active streamers. METAL collects this data by building the best retroactive clipping software in the world. In other words, you don't need to be consciously recording. You actually just have METAL on in the background while you're playing, and you hit a button to clip the last 30 seconds after something interesting happens. It's very similar to how Tesla and self-driving does bug reporting, if you have ever done a self-driving bug report in Teslas. The result is that METAL has accumulated 3.8 billion clips of the best moments and actions in games, resulting in one of the most unique and diverse datasets of peak human behavior actively mining for the interesting moments. They were also very prescient in navigating privacy and data collection concerns by mapping actions to these visual inputs and game outcomes. As you saw on our Fei-Fei Li and Justin Johnson episode with World Labs, and with the recent departure of Yann LeCun from Meta, there's a lot of interest in world models as the next frontier after LLMs, to improve on spatial intelligence and to work on embodied robotics use cases. DeepMind has been working on this with Genie 1, 2 and 3 and CIMA 1 and 2 And this year, OpenAI seem to finally agree because they have been betting on LLMs a lot. And they made the news by offering $500 million for Medal's video game ClipData. Our guest today, Pim, turned down that money and instead chose to build an independent world model lab instead. Khosla Ventures led the $134 million seed round, which is Vinod Khosla's largest single seed bet since OpenAI.
We were able to get an exclusive preview of GI's models, which unfortunately we cannot show you directly. But I can confirm, they were incredibly human-like and we chose to include the first 11 minutes of the demo discussion even though I couldn't show it to you. It may be hard to follow, but I tried to call out what was noteworthy for you to know as your likely reaction if you were watching along with us. Now enjoy the world's first look at my first look at General Intuition.
**Pim** (2:08)
So what I'm about to show you is a completely vision-based agent that's just seeing pixels and predicting actions the exact same way a human would. And so yeah, what I'll show you here is what this looks like four months ago. So this was, so again, this is just an agent that's seeing, that's receiving frames and it's just predicting action. So you can see it has like a decent sense of being able to navigate around. It tabs a scoreboard just like gamers always tab the scoreboard. So these are purely, these are pure imitation learning.
**Unknown Host** (2:39)
I see, so this is slicing a knife.
**Pim** (2:41)
Yeah, exactly. So it's doing everything like humans would in this case. Here was the first interesting part that we saw, like it gets stuck and then it has, they have memory as well. So you see, it can get unstuck. How long is the memory? Four seconds. Yeah, four seconds for the straights and colors. Okay, so this was four months ago. This was maybe a few weeks after that. So you can see there is like, it's still doing the scoreboard thing, but it's, they're still, they're still quite like, and these are bots too. So you can see it.
**Unknown Host** (3:06)
It's very human. Let's just say that.
**Pim** (3:07)
Yeah. And then, right, so this was really like the early days of research, where you can see right, there's one thing, and then goes for another. And then we've been scaling, right, on data and compute, and also we've just been making the models better. And this is where we are now. So what you're seeing is pure, like I said, pure mutation learning. This is just the base model. There's no RL, no fine-tuning. This model sees no game states. It is purely capable, not in sequence, in sentence. It's purely predicting the actions from the frames. That's it. And this is playing against real humans, just like a human would play. And it's also, it's running completely in real time. So there's absolutely, everything here plays exactly like human.
61 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000748427731