SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow) artwork

SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)

Latent Space: The AI Engineer Podcast

December 18, 2025

As with all demo-heavy and especially vision AI podcasts, we encourage watching along on our YouTube (and tossing us an upvote/subscribe if you like!
Speakers: Swix, Joseph Nelson, Pengchuan Zhang, Nikhila Ravi
**Swix** (0:03)
Okay, we're here in the remote studio with the grand return of the Roboflow and Latent Space and Sam combo. Welcome to Joseph, my sort of vision co-host, I guess.

**Joseph Nelson** (0:15)
Thanks. Good to be here.

**Swix** (0:16)
Welcome back. We also have, welcome back, Nikhila Ravi, who's the lead on SAM 2 I guess just SAM in general, right? And we have, joining us, Pengchuan, who's also a researcher on SAM.

**Pengchuan Zhang** (0:27)
Yeah, nice to meet you guys.

**Swix** (0:29)
So congrats on SAM 3's launch. I mean, like the demo, each time you set it up, like really amazingly. And I think like every time my general impression or takeaway when I tell people about SAM is like, just the every time you have a new release, like it's like once a year you show up, you drop a banger and then you like, you know, you just like drop the mic and go for next year. And you also add a dimension. So I was entirely, like, weirdly not surprised when SAM 3 had the 3D thing, because I'm like, well, yeah, which is the next dimension to go? It's like 3D.

**Nikhila Ravi** (1:02)
Yeah, actually, maybe just on that, I think that's actually a common misconception. We launched directly three separate models this time. It was SAM 3, SAM 3D Objects and SAM 3D Body.

**Swix** (1:16)
Yes.

**Nikhila Ravi** (1:16)
Those were two completely separate models and SAM 3 is just the image and video understanding model.

**Swix** (1:23)
Which is on a deader backbone and is sped up. Yeah, sorry, I didn't mean to sort of pre preface all this. But maybe for just to remind our audience or maybe for people in new to the SAM series of a podcast that we've done so far, maybe each of you can sort of go around and intro like your, or your sort of entry into computer vision or sort of your relationship with SAM. Go ahead, Nikki.

**Nikhila Ravi** (1:45)
Okay, cool. Hi, everyone. I'm Nikhila. I'm a researcher at Meta. I've been at Meta for eight and a half years. So I really been through evolution of the field in that time. I really started working on a range of different problems in computer vision, worked briefly on 3D. We've got this library called PyTorch 3D. But really started on this Segment Anything as a project in around sort of late 2021 So it's actually been almost four years since I've been like working on this Segment Anything space. And we started with Sam 1 in 2023, Sam 2 last year in July 2024, and then now Sam 3 So it's been a combination of a lot of work of a lot of people over the years. So yeah, really, really excited to be at this point and get to share it with all of you. I'll hand it over to Pengchuan.

**Pengchuan Zhang** (2:43)
Yeah. Hello, everyone. So I'm Pengchuan. I'm a researcher at Santin. I have been working in computer vision this field for nearly nine years, starting from 2017 I think it's a long time. I have been working in MSR for five years, and then kind of moved to Meta Reality Lab to work on egocentric foundation models on AI glasses for a while. And then in 2023, way near, I moved to SAM team, and that time is exactly the start time of SAM SLUY. And really, I think that's the lifetime experience I have on the SAM SLUY team. And it's glad that SAM SLUY is out, and I kind of achieved my original grand goal of computer vision to reach kind of human performance of detection, segmentation, tracking, image and videos.

**Joseph Nelson** (3:32)
I'm Joseph, co-founder, CEO at Roboflow, where our mission is to make the world programmable. We think software should have the sense of sight and models like SAM and others are critical to unlocking that capability. Now millions of developers have the Fortune 100, build with Roboflow's tools and infrastructure to create and deploy models to production. We've been big believers of the Meta family of open source models all the way back to like Mask R, CNN and Detectron 2, all the way to presence of SAM 1, SAM 2 and SAM 3 The work that the Meta team does to advance state of the art in open source computer vision has been bedrock to enabling developers and enterprises globally to adopt AI. So we've been big fans of the work and I'm pleased to be joining you today, Swix, to co-host the episode on SAM 3

**Swix** (4:21)
And you guys shipped your own DETTER model too.

**Joseph Nelson** (4:24)
Yeah, we've been doing some work to advance machine learning research too. Like one of the, for example, DETTER Detection Transformers, which was born out of NeurIPS last year. I think, Swix, you actually challenged us. You were like, hey, what are some of the advancements that are happening in computer vision and in visual AI? And we had this observation that Transformers had surpassed a lot of CNNs in vision tasks, but they hadn't been made to run real-time, as in, you know, over 30 frames per second, for example, on like a small T4, or excuse me, small like edge device, and hundreds of frames per second on like a T4. We did some research and published RF data Roboflow Detection Transformer, which is, you know, we kind of joked the greatest of all time model for doing real-time segmentation and object detection on the edge. Now in RF data, it's, you know, you have to have a fixed class list and need to know some of the objects that you want to segment at a time. But for anyone that's running on like constrained compute and on an edge device and wants like an Apache 2 model to do that, RF data and its family of models are key to fulfilling that mission and that goal.

61 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000748427730