**Div Garg** (0:00)
I think OpenAI has definitely lost a lot of the lead they used to have. The GPD 4, where they were kind of the sole winner and no one could catch up with them. At this point, it does seem like a very homogenous market where everyone is kind of close. When anything is disruptive, I think it takes a lot of time for that to catch on. And I think we are at the start of the disruptive era, where a lot of the online communication and interaction will get disrupted by agents. Humans are pretty good at navigating websites with UI. And so theoretically, AI can also become very, very good, even better. I would call it an explosion of applications, which I don't think has happened so far. If you think about agent applications, I think it's still very good.
**Nathan Labenz** (0:40)
Hello, and welcome back to The Cognitive Revolution. Today, Div Garg, founder and CEO of MultiOn, returns for his third appearance on the show. A lot has changed in the AI agent landscape since I first spoke with Div in mid-2023. As you might remember at that time, the AI community was abuzz about the potential for AI agents, with projects like BabyAGI giving large language models access to tools and really only very minimal guidance, and then stepping back to watch and see what they could accomplish on their own. Of course, as it turned out, they didn't accomplish all that much. While there were many amazing moments, those were the outliers. On average, the step-by-step error rate, including on very mundane microtasks, was too high for agent frameworks to successfully string many-step sequences together all that often, and we also learned that while large language models can improve in many areas through self-critique, they have a tendency to get stuck on obstacles that humans quickly find ways around. For that reason, much of the last 18 months of work on agents has gone into developing better and more prescriptive scaffolding, with many companies ultimately delivering platforms for what I call intelligent workflows. That is, workflows that a human has designed, and where the AI is needed to do some important subtask, which requires intelligence, but where the AI is not given freedom to choose its own adventure. As of January this year, Div and the MultiOn team were still among the most bullish on open-ended agents, and as you'll hear in this conversation, they have continued at least partially to buck that trend. They have built some new scaffolding and they have developed interesting techniques for domain-specific fine-tuning, but their agent continues to take arbitrary natural language requests and gamely does its best to fulfill them. The progress I found in my testing is pretty obvious, and in some contexts the company claims human-level performance, but still the system as a whole is not a viable substitute for a human assistant. With that in mind, I was excited to pepper-tip with questions about what he's learned from all of this activity. And so in this conversation, we unpack the latest in agent development, including the company's data collection strategy, the seemingly missing market for human computer use data, and the role of synthetic data in bridging that gap. The company's model strategy, including what models they've chosen as base, what fine-tuning techniques they're using, and how their computer vision approaches have evolved over time. Why benchmarks so often show human-level performance while the real-world results are clearly not as strong? The future of agent authentication, as well as which parts of the internet at large will compete to serve agents versus which parts will try to exclude them. And finally, what sorts of customers MultiOn is looking to partner with now, as well as how they're thinking about competing with hyperscalers in light of Cod's new computer use capability.
Overall, it's clear to me that while it's taken longer than I had expected, reliable agents that can perform a very large percentage of routine computer use tasks are coming. It's only a matter of time. And as you'll hear, Div agrees with recent suggestions from both OpenAI and Anthropic that 2025 will be the year. That of course makes Div a very busy man, and so I very much appreciated his time and how open he was willing to be about the path that MultiOn has taken and the lessons they've learned along the way. As always, if you're finding value in the show, we'd appreciate it if you'd share it online, write a review on your podcast app, or leave a comment on YouTube. Of course, we welcome your feedback and suggestions via our website, cognitiverevolution.ai, or by DMing me on your favorite social network. Now, here's my conversation with Div Garg of MultiOn, catching up on the last year in AI Agents. Div Garg, founder and CEO of MultiOn, welcome back to The Cognitive Revolution.
84 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000679069338