Game Controllers Are the Universal Robot Interface | Pim de Witte artwork

Game Controllers Are the Universal Robot Interface | Pim de Witte

MTS

October 7, 2026

General Intuition Co-Founder and CEO Pim de Witte discusses their $220M Series A raise at a $6.2B valuation, proving sim-to-real transfer to real-world robotics using game controllers, and their thesis on scaling world models natively in pixel space. Turn ideas into software people love.

Speakers Pim de Witte, Sophia

TopicsNews

Pim de Witte (0:00)

The beauty of this approach is that game controllers have already built this general interface that converts human intuition or human intelligence into these very general outputs that, by the way, almost every robot in the world ships with, right? And so I think it's just very practical. And I think our bet is just like that happened with code, it will happen to pixels. And we are just sort of the lab that's scaling pixels all the way and entropic skills code all the way.

Sophia (0:24)

Hello everyone, and welcome to MTS. Today, I am joined by Pim de Witte. He is the co-founder and CEO of General Intuition. They're building an AI lab that builds models that learn how to act from gameplay and other action labeled video. And last week, they just closed another $220 million at a $6.2 billion valuation. So Pim, welcome to MTS.

Pim de Witte (0:50)

Thank you. I appreciate you inviting me.

Sophia (0:52)

Yeah, welcome to the show.

One thing I wanted to ask about is this recent round comes pretty recently since your last round. I know you guys recently raised, not even that long ago, another giant round. And then you just announced you had a $320 million Series A. This is like about three months before. And so, yeah, I just wanted to ask a little bit about this new funding round. What do you guys hope to do? Why raise again? And what's leading to this big growth?

Pim de Witte (1:21)

The big change that happened in between the two rounds is that we successfully proved transfer to real world robotics of the models. So one way of thinking about it is humans can play video games using the same primitive. So visions or their eyes and then outputs, actions as they can tell operate robots. So people write, people do lots of tell operations to request three. So they can tell operate using keyboard and mouse. They can tell operate using game controllers. And so what we were able to prove is that just like people can tell operate and play video games using the same intelligence that our models are able to operate in a similar way, where they can both play video games at the highest level off just pixels, but also control real world robots. And so before we raised the last round, we had not yet proven the recipe out in simulation and games. We had not yet successfully transferred it over to real world robots. And as we scaled data and compute and we got better at building models, we were able to prove that. And I think that was the reason why a lot of new capital came into the company.

Sophia (2:30)

What would you say right now is one of the hardest parts about building these world models?

Pim de Witte (2:39)

I think it's unclear how it plays with frontier LLMs and text models and sort of how everything converges. I think there's a very large push for coding. And then simultaneously, we're making a very large push for pixel space. So think of coding as one way of going about the problem, where you see really impressive demos of Astra or Claude controlling robots. But that's using code, which means it's usually quite slow because that code has to be written, often compiled, deployed onto a robot.

And so we sort of build models that are natively in pixel space, and they can be a lot faster. But I think a lot of the question marks that I have, and I think that we're exploring is how do these things come together? How do you get both the best of coding, but also the best of these vision-based policies that we're working on? And world models here is a broad term. I just view world models as generally natively frame-based or pixel-based models.

And it's going to be an exciting few months to see how it all plays out. I think everyone in the space ships on models that people love, and I think we're very close to doing that. You can expect to see some models from us this year.

Sophia (3:49)

One thing I was curious about is the actual name General Intuition. And if you think of that, it's like, how do you build AI in a way that it has some of this general intuition type of things rather than just depending solely on just words? So tell me a little bit about the origin name here and how it connects to the company.

Pim de Witte (4:06)

So it's actually named after Demis Hasavis, and it's less immediately actionable. It's more of a north star for us. So when the DeepMind team tried to do Alpha Fold, they originally tried to train an AI bot, a computer use agent, if you will, to play Foldit to basically generate these protein sequences inside a physics engine. You can read on Foldit, it was a video game that lots of people played.

16 more minutes of transcript below

Thousands of transcripts fetched by people building searchable podcast archives

Fetch the whole transcript

The demo key returns a sample episode in full, no card needed:

request
curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Markdown with the speakers named, for your notes, your knowledge base, or anything that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.

Using your own key:

request
curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000793628047