Bots Already Make Up 60% of All Web Traffic | Kyle Jeong
MTS
September 29, 2026
BrowserBase growth engineer Kyle Jeong breaks down how computer use agents evolved from slow, inaccurate coordinate clicking to executing code and leveraging accessibility tree snapshots.
Speakers Kyle Jeong
TopicsNews
Kyle Jeong (0:00)
Yeah, I think there's a big gap in trust and authentication, and then also like just access. You've probably seen like Muse or stuff from Meta, and they have like an ongoing battle with Amazon. Amazon doesn't want Muse to be able to go and visit Amazon and purchase things or put them in the cart. A lot of the internet, they don't want just any bots rolling rampant onto their website. I think the tricky part here is how do we broker this verified version of like, how do I say that my agent is a good agent?
SPEAKER_2 (0:32)
And we're back. I'm here with Kyle Jeong from BrowserBase. He's a growth engineer.
BrowserBase builds web browsers for AI agents. He's worked on Stagehand and the BrowserBase MCP, was a Kleiner Perkins Fellow and engineer at Pace. So welcome Kyle, great to have you.
Kyle Jeong (0:48)
Yeah, glad to be on.
SPEAKER_2 (0:50)
So I have a lot of questions, especially about computer use and how it works. So I wanted to ask you about your like sub stack article, where you talk about how does Astra's computer use actually work. It's something that I've already been like very blown away by, because I feel like, you know, even like the Cloud computer use pales in comparison.
So yeah, I want to hear your takes.
Kyle Jeong (1:12)
Yeah, yeah, so I mean, we can talk about, I guess, like a short history of the computer use.
So before, I guess, like the first versions from like Anthropic and even Operator from OpenAI, it was all like coordinate click based. So you'd give the model a screenshot, you give it like a goal, a task, and it would return like, okay, if I want to click this button, I'm going to return like click, like that's the action, and then exactly the coordinates of like where that button is on a page. And so it worked like kind of okay, like it was good in demos. And so we were like able to spin up like versions of operator that worked, but it was really slow. And it was kind of like inaccurate since the models were kind of trained to only know like a certain viewport size. So like, let's say we've found one of them that's common is like 1288 by 711, which is kind of like a weird screen size. But if you give it an environment or a computer that's not the exact size, it just kind of misses all the clicks. So that's kind of like an earlier version. But over time, we kind of just saw it get better all of a sudden. And what these new models are doing, like Astra and even the models from Anthropic, is the models are kind of just really good at writing code. Obviously, coding is verifiable and so they got really good at coding really fast. But what if, let's say we take code and this idea of just writing code to execute tools, but also like control the browser or computer, what if we just let them write code? And so models got really good at writing. Playwrights is an example of a framework that models are really good at writing. And so Astra specifically is able to take a screenshot or also like a snapshot of a page, which I'll talk about in a sec, but it takes a snapshot screenshot and then can write code to, let's say, click the button or type this in. And so instead of trying to find the coordinates on the page, and the exact like pixel XY coordinates, it just writes code to interact. And so it just got way faster in one and then it's more repeatable.
The way that Astra works is they have like a node ripple that saves actions that might be repeated so that it doesn't have to generate code multiple times over again. Right.
And so yeah, we've kind of just seen like models get good at writing code. They get good at like executing code to control the browser or computer. And so that's kind of what they do. And then I guess like to talk about the snapshot.
Screenshot, obviously everyone knows what a screenshot is. You can take a screenshot on your laptop and you'd see like the images that's on the page. A snapshot is kind of different in the sense that it takes what's called accessibility tree. So it's meant for like if you're maybe like visually impaired to be able to interact with your phone or your web browser or even your computer. And so actually every website that you can like open in a Chrome browser will have an accessibility tree. And this will remove like, okay, if you don't see the page, you don't need to see the visual, like the style. So remove all the CSS. We don't need to have like that kind of visual context, but we have like a representation of what's on the page and then what actions you can take. And so that's actually like more helpful context to the models. So screenshot, snapshot, that's kind of what goes in the difference.
16 more minutes of transcript below
Thousands of transcripts fetched by people building searchable podcast archives
Fetch the whole transcript
The demo key returns a sample episode in full, no card needed:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090Markdown with the speakers named, for your notes, your knowledge base, or anything that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000792201402