**Gergely Orosz** (0:00)
What happens when you let AI agents ship code for months, and no developer reads a single line? Today's guest tried exactly that. He built a light-off software factory, and four months later, he had no choice but to shut it down as things just stopped working.
Dex Horthy is the founder of HumanLayer, and the person who coined the term context engineering days before Andrei Karpathy and Tobi Lutske made it famous. He spent the last two years talking to hundreds of AI engineers about what actually works when you build with LLMs, and is testing the most succinct ideas with his own team. In today's conversation, we will discuss Context Engineering, what it is, and the physics of context windows, including what the dump zone is. Loop engineering, from the Ralph Wilkamp technique to the slow loop that Dex's team runs every night to wake up to code clean up PRs. The rise of software factories, from a NATO conference in 1968 through DevOps to today's agentic factories. Spec driven development and why specs always drift from the code itself. And many more. If you want to understand increasingly important concepts like concept engineering and harness engineering, or want to know how far you can push the let agents build everything idea, from someone who pushed it further than almost anyone, then this episode is for you. This episode is presented by Antisysys. If you work with agents, your job is no longer just writing code, it's specifying and testing it. And Antisysys is the most effective method of verifying agentic code today.
Today's episode is brought to you by Buildkite, the CIO certification platform trusted by OpenAI, Entropic, Cursor, Nvidia, Uber, Canva and more. Today we're talking about pushing the right context into models so that they write better code. Right after that starts working, your agents will write more code, a lot more. Trusting that code avalanche is where many teams face a challenge today. Every change that an agent makes still has to be built, tested and proven safe before it ships.
Worked on my machine is not enough, so you obviously need CI. But when agents are pushing 5, 10 or 50 times the commit volume to your pipelines, faster CI runners won't save you. Shaving 30 seconds off a single build is meaningless when the queue is 100 plus jobs deep. What you really want is a CI system that gets faster as the volume grows, and CI that offers instant parallelization to give you unlimited concurrency and to intelligently route changes at runtime. This is what Buildkite does and why global software leaders continue to rely on it. The same architecture that observed the scale of Shopify and Uber a decade ago now runs about 1.4 billion job minutes a week across cursor, meta, Reddit and Snowflake. While the rest of the CI world are cracking under the weight of re-architecting their platform, Buildkite continues to reliably grow. Agents running on your infrastructure are Buildkite. Any cloud, any chip, your secrets, your scale. Every artifact and log is captured, so when something fails, either you or your agent have immediate insight for why. As you're entering the context you'll give to your agents, think about how you'll verify what they hand back. If your system is buckling under the increased volume, head to buildkite.com/pragmatic. 30-day all-access trial, no credit card, and an actual human engineer on standby. His name is Ola, and he's very helpful.
So Dex, welcome to the podcast.
**Dex Horthy** (2:51)
Super stoked to be here, dude.
**Gergely Orosz** (2:52)
Before we get into some of the context engineering and some of the more spicy stuff as well, how did you get into tech? How did you follow up with computers?
**Dex Horthy** (3:01)
Oh man, so I was doing undergrad as a physics major, and I realized that I didn't like academia. And there's basically like two or three paths out of physics is basically you go get a PhD, or you go into finance, or you go do programming. At that time, this was 2012, 2011, when I was in the middle of undergrad and deciding what to do. And I had done an internship when I was in high school, I was working with NASA researchers to a jet propulsion lab in California. They had just gotten this really high fidelity, the most fine-grained data set of altitudes, like the heights of topographical map of the south pole of the moon.
And the south pole of the moon is really interesting because some of the craters there are so deep. Because of the angle it has, it got hit by meteor storms like no other part of the moon. So there's very deep craters that have never seen sunlight. And so there's frozen liquid water in there from the formation of the moon. And so scientists were really interested in getting down there and exploring. And so we had this really fine-grained map and it was like, okay, cool, let's build software so that I have point A to point B. I know the limitations of my rover. Can max incline up as this, max incline down as that. Find a path from point A to point B that doesn't break those rules of incline. So I was 17 I had never cracked a CS textbook. So I wrote, I basically wrote a really naive, bad version of Dijkstra's algorithm for pathfinding.
101 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000776932678