**Nathan Labenz** (0:02)
Before we dive into today's episode, I want to tell you about a new show from Turpentine called Modern Relationships. On the season ahead, I sit down with power couples in tech and leading relationship thinkers to explore how ambitious people actually make partnerships work. Whether you're dating in a relationship or just curious how technology is reshaping modern love, I think you'd enjoy this on your feed. Our first episode features founders funds Delian Asparuhov and tech researcher Nadya Asparuhov, who take us through their evolution from dating to marriage to parenthood, with absolutely no filter on the challenges and growth along the way. You can find Modern Relationships wherever you get your podcasts. Now, on to today's episode.
**Will Hardman** (0:37)
Is multimodal understanding in an AI important on the path towards AGI? It's not entirely clear that it is, but some people argue that it is.
So one reason that one might want to research these things is to see if by integrating the information from different modalities, you obtain another kind of transformational leap in the ability of a system to understand the world and to reason about it. I would say in Inverted Commons, similarly to the way we do. For open-source researchers, the last few months have really seen the arrival of these huge interleaved datasets which has really jumped the pre-training dataset size that's available. I'm amazed at the Perceiver Resampler works because it feels to me just like tipping the image into a blender, pressing on, and then somehow when it's finished training, the important features are retained and still there for you.
**Erik Torenberg** (1:33)
Hello, Happy New Year, and welcome back to The Cognitive Revolution. Today, I'm excited to share an in-depth technical survey covering just about everything you need to know about vision language models, and by extension, how multimodality in AI systems currently tends to work in general. My guest, Will Hardman, is founder of AI advisory firm Veritai, and he's produced an exceptionally detailed overview of how VLMs have evolved, from early vision transformers to CLPS pioneering alignment work, to today's state-of-the-art architectures like InternVL and Llama3V. We'll examine key architectural decisions like the choice and trade-offs between cross-attention and self-attention approaches, techniques for handling high-resolution images and documents, and how evaluation frameworks like MMMU and Blink are revealing both the remarkable progress and the remaining limitations in these systems. Along the way, we dig deep into the technical innovations that have driven progress, from Flamingo's Perceiver Resampler, which reduces the number of visual tokens, to a fixed dimensionality for efficient cross-attention, to InternVL's dynamic high-resolution strategy that segments images into 448x448 tiles while still maintaining global context. We also explore how different teams have approached instruction tuning, from Lava's synthetic data generation to the multi-stage pre-training approach pioneered by the Chinese research team behind QuenVL.
Our hope is that this episode gives anyone who isn't already deep in the VLM literature a much better understanding of both how these models work and also how to apply them effectively in the context of application development. Will spent an estimated 40 hours preparing for this episode, and his detailed outline, which is available in the show notes, is probably the most comprehensive reference we've ever shared on this feed. While I have not worked personally with Will outside of the creation of this podcast, the technical depth and attention to detail that he demonstrated in what for him is an extracurricular project was truly outstanding. So if you're looking for AI advisory services and you want someone who truly understands the technology in-depth on its own terms, I would definitely encourage you to check out Will and the team at Veritai.
Looking ahead, I would love to do more of these in-depth technical surveys, but I really need partners to make them great. There are so many crucial areas that deserve this kind of treatment, and I just don't have time to go as far in-depth as I'd need to do them on my own. A few topic areas that are of particular interest to me right now include, first, recent advances in distributed training. These could democratize access to frontier model development, but also pose fundamental challenges to compute-based governance schemes. Next, what should we make of the recent progress from the Chinese AI ecosystem? Are they catching up by training on Western Model Outputs, or are they developing truly novel capabilities of their own? There's not a strong consensus here, but there's arguably no question more important for US policy makers as we enter 2025 I'm also really interested in biological inspirations for neural network architectures, or any comparative analysis of human and artificial neural network characteristics. The episode that we did with AE Studio stands out as one of my favorites of 2024, and I would love to have a more comprehensive understanding of what we collectively know about this space. I'm similarly interested in the state of the art when it comes to using language models as judge or otherwise evaluating model performance on tasks where there's no single right answer. This is a problem that we face daily at Waymark, and which seems likely to have important implications for how well reinforcement learning approaches will work and scale in hard to evaluate domains. Finally, for now, I would love to catch up on the latest advances in vector databases and rag architectures. I've honestly been somewhat disillusioned with embedding-based rag strategies recently, and I've been recommending Flash Everything as the default relevance filtering strategy for a while now. But I do wonder, what might I be missing? In any case, the success of our previous Survey Style episodes, including our AI Revolution in Biology episodes with Amelie Schreiber, and our Data Data Everywhere Enough for AGI episode with Nick Gannon, suggest that people find these detailed overviews to be a helpful way to catch up on important AI subfields. So, if you have or are keen to develop deep expertise in an area that you think our audience would benefit from understanding better, please do reach out. I'm open-minded about possible topics and very interested to hear what you might propose. You can contact us, as always, via our website, cognitiverevolution.ai, or by DMing me on your favorite social network. And of course, I'm more than happy to give you the chance to plug your product or services as part of your appearance on the show. Now, I hope you enjoy my conversation with Will Hardman, AI Advisor at Veritai, about all aspects of Vision Language Models. Will Hardman, AI Advisor at Veritai, and AI Scout on all things Vision Language Models. Welcome to The Cognitive Revolution.
191 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000682552612