**Charlie** (2:11)
Happy August. This is your AI co-host, Charlie. Thank you for all the love for our special 1 million downloads Winds of AI Winter episode last week, especially Sam, Archie, Trellis, Morgan, Shrey, Han and more. For this episode, we have to go all the way back to the first viral episode of the podcast, Segment Anything Model and the Hard Problems of Computer Vision, which we discussed with Joseph Nelson of Roboflow. Since Meta released Sam 2 last week, we are delighted to welcome Joseph back as our fourth guest co-host to chat with Nikhila Ravi, Research Engineering Manager at Facebook AI Research and lead author of Sam 2 Just like our Sam 1 podcast, this is a multimodal pod because of the vision element, so we definitely encourage you to hop over to our YouTube, at least for the demos, if not our faces. Watch out and take care.
**swyx** (3:10)
Welcome to the Latent Space Podcast. I'm delighted to do Segment Anything 2 One of our very first viral podcasts was Segment Anything 1 with Joseph. Welcome back.
**Joseph Nelson** (3:19)
Thanks so much.
**swyx** (3:20)
Then this time we are joined by the lead author of Segment Anything 2, Nikki Ravi. Welcome.
**Nikhila Ravi** (3:25)
Thank you. Thanks for having me.
**swyx** (3:26)
There's a whole story that we can refer people back to episode of the podcast way back when, for the story of Segment Anything. But I think we're interested in just introducing you as a researcher on the human side. What was your path into AI research? Why did you choose computer vision coming out of your specialization at Cambridge?
**Nikhila Ravi** (3:46)
Sure. I did my undergraduate degree in engineering at Cambridge University. The engineering program is very general. First couple of years, you study everything from mechanical engineering to fluid mechanics, structural mechanics, material science, and also computer science. Towards the end of my degree, I started taking more classes in machine learning and computational neuroscience, and I really enjoyed it. Actually, after graduating from undergrad, I had a place at Oxford to study medicine. I was initially planning on becoming a doctor, had everything planned, and then decided to take a gap year after finishing undergrad. Actually, that was around the time that deep learning was emerging, and in my machine learning class in undergrad, I remember one day our professor came in, and that was when Google acquired DeepMind. That became a huge thing. We talked about it over the whole class. It really kicked off thinking about, okay, maybe I want to try something different other than medicine. Maybe this is a different path I want to take. Then in the gap year, I did a bunch of coding, worked on a number of projects, did some freelance contracting work, and then I got a scholarship to come and study in America. I went to Harvard for a year, took a bunch of computer science classes at Harvard and MIT, worked on a number of AI projects, especially in computer vision. I really enjoyed working in computer vision, applied to Facebook and got this job at Facebook, and I'm now at Facebook at the time, now Matter, and I've been here for seven years.
Very circuitous path, probably not a very unconventional. I didn't do a PhD. I'm not a typical research scientist, definitely came from more of an engineering background. But since being at Matter, I've had amazing opportunities to work across so many different interesting problems in computer vision from 3D computer vision. How can you go from images of objects to 3D structures, and then going back to 2D computer vision and actually understanding the objects and the pixels and the images themselves? It's been a very interesting journey over the past seven years.
**swyx** (6:05)
It's weird because I guess with Segment Anything 2, it's like 4D because you solve time. You started with 3D and now you're solving the 4D.
**Nikhila Ravi** (6:14)
Yeah, it's just going from 3D to images to video. It's really covering the full spectrum. Actually, one of the nice things has been, so I think I mentioned I wanted to become a doctor, but actually Sam is having so much impact in medicine, probably more than I could have ever had as a doctor myself. So I think hopefully Sam 2 can also have similar impact in medicine and other fields.
**swyx** (6:39)
Yeah. I want to give Joseph a chance to comment. Lizette, we know your story about going into vision, but in the past year since we did our podcast on Sam, what's been the impact that you've seen?
**Joseph Nelson** (6:51)
Segment Anything set a new standard in computer vision. Recapping from the first release to present. Sam introduces the ability for models to near zero-shot, meaning without any training, identify perfect polygons and outlines of items and objects inside images. That capability previously required lots of manual labeling, lots of manual preparation, clicking very meticulously to create outlines of individuals and people. And there were some models that attempted to do zero-shot segmentation of items inside images, though none were as high quality as Segment Anything. And with the introduction of Segment Anything, you can pass an image with Sam 1, Sam 2 videos as well, and get pixel-perfect outlines of most everything inside the images. Now, there are some edge cases across domains, and similar to the human eye, sometimes you just say which item you maybe you most care about for the downstream task and problem you're working on. Though Sam has accelerated the rate at which developers are able to use computer vision in production applications. At Roboflow, we were very quick to enable the community of computer vision developers and engineers to use Sam and apply it to their problems. The principal ways is using Sam. You could use Sam as is to pass an image and receive back masks. Another use case for Sam is in preparation of data for other types of problems. For example, in the medical domain, let's say that you're working on a problem where you have a bunch of images from a wet lab experiment. From each of those images, you need to count the presence of a particular protein that reacts to some experiments. To count all the individual protein reactions, you can go in and lab assistants to this day will still individually count and say, what are the presence of all of those proteins? With Segment Anything, it's able to identify all of those individual items correctly. But often, you may need to also add a class name to what the protein is, or you may need to say, hey, I care about the protein portion of this, I don't care about the rest of the portion of this image.
48 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000664657872