The Computer Vision Revolution with Junnan Li and Dongxu Li of BLIP and BLIP2 artwork

The Computer Vision Revolution with Junnan Li and Dongxu Li of BLIP and BLIP2

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

March 9, 2023

As recently as January 2021, the challenge of "interpreting what is going on in a photograph" was considered "nowhere near solved.
Speakers: Erik Torenberg, Nathan Labenz, Junnan Li, Dongxu Li
**Erik Torenberg** (0:00)
Turpentine is a network of podcasts, newsletters, and more, covering tech, business, and culture, all from the perspective of industry insiders and experts.
We're the network behind the show you're listening to right now.
At Turpentine, we're building the first media outlet for tech people by tech people. We have a slate of hit shows across a range of topics and industries, from AI with Cognitive Revolution to Econ 102 with Noah Smith. Our other shows drive the conversation in tech with the most interesting thinkers, founders, and investors, like Moment of Zen and my show Upstream. We're looking for industry-leading hosts and shows along with sponsors. If you think that might be you or your company, email me at erik.turpentine.co. That's E-R-I-K at turpentine.co.

**Nathan Labenz** (0:45)
In a way, it's not that dissimilar from how we see, right? Like we have our eyes, the eyes kind of take in raw light and turn that into a signal, and that signal goes through the nerve and finally gets back to the back of the brain. And by that point, it's not that interpretable either, right? It doesn't necessarily correspond to language. But then there's some further connector that like turns that visual data into something that I can understand as language or at least understand and then articulate as language.
So it feels like there is something kind of analogous taking shape in the AI world.

**Junnan Li** (1:22)
Imagine you are human, you grow up learning only knowledge. And now one day you open your eyes, you don't really know how to interpret what you see. So that's what we are going to do here to kind of build this bridge between these two modalities.

**Nathan Labenz** (1:40)
Hello and welcome to The Cognitive Revolution, where we interview visionary researchers, entrepreneurs, and builders working on the frontier of artificial intelligence.
Each week we'll explore their revolutionary ideas, and together we'll build a picture of how AI technology will transform work, life, and society in the coming years. I'm Nathan Labenz, joined by my co-host Eric Thornburg.

**Erik Torenberg** (2:03)
Omneky uses generative AI to enable you to launch hundreds of thousands of ad iterations that actually work, customized across all platforms with a click of a button. I believe in Omneky so much that I invested in it, and I recommend you use it too.
Use Cograv to get a 10% discount.

**Nathan Labenz** (2:21)
Today's episode was a fun one for me. Researchers Junnan Li and Dongxu Li, both of Salesforce Research's Singapore office, have co-authored some of the most practically useful computer vision papers of the last year. As recently as January 2021, the challenge of using AI to interpret what is going on in a photograph was considered to be nowhere near solved. But just a year later, Junnan and Dongxu changed all that by publishing and open sourcing Blip, a family of pre-trained models that delivered state-of-the-art performance on image captioning, visual question answering, and image text matching. For Waymark, my company, Blip was a godsend. Suddenly, we had a reliable way to understand the contents of users' images, allowing us to make useful image suggestions for the very first time. This was something we had worked toward for years.
Unusually in today's AI landscape, Blip has held the title of best image captioner for over a year, ultimately becoming the 18th most cited AI paper of 2022 And more recently, just as worthy rivals to Blip started to come online, Junnan and Dongxu changed the game again with Blip2. As an aside, for a funny moment in Cognitive Revolution history, you can listen back to episode 1, in which Suhail tells me about the release of Blip2 live on the show, forcing me to clear my calendar for the rest of the afternoon to go check it out.
Now, Blip2 uses a different approach, which I think may ultimately prove even more influential. Rather than training a large model end-to-end, this time they trained a much smaller model that connects a frozen vision model to a frozen language model. This strategy has several benefits. First, because it injects semantic visual information into the language model's latent space, you can now have an open-ended dialogue about an image in which the language model shows remarkably detailed and nuanced understanding.
Second, because the connector model is so much smaller, training time and cost are dramatically reduced. Blip2 was trained on just a single A100 machine in less than 10 days, making it easy to upgrade the system as new and more powerful language models become available. You can even use your own fine-tuned language models as well.
Small connector models like Blip2 show just how much potential remains to be drawn out of today's large language models and seem likely to play an important role in the great implementation of multimodal AI across society. One note for listeners, both Junnan and Dongxu are native Chinese speakers, and while both are perfectly fluent in English, audience members listening at 2x speed might benefit from watching this episode on YouTube, where we've also included subtitles for your convenience. Now enjoy our conversation with Junnan Li and Dongxu Li. Junnan Li and Dongxu Li, welcome to The Cognitive Revolution.

57 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000603434116