The cost of intelligence will never be this cheap again, the failure of intensive specs, and how bots disguise inefficient workflows artwork

The cost of intelligence will never be this cheap again, the failure of intensive specs, and how bots disguise inefficient workflows

Dev Interrupted

May 29, 2026

Are we officially entering the "Eternal Sloptember"? This week on the Friday Deploy, Ben and Andrew unpack the quiet rebellion against skyrocketing API costs as teams transition to fine-tuned local models.
Speakers: Ben Lloyd-Pearson, Andrew Ziegler

Topics: Technology

**Ben Lloyd-Pearson** (0:05)
Opus 4.8 drops, right, as we come on to record this show. Have you sent it any requests yet, any prompts?

**Andrew Ziegler** (0:12)
Not yet, but I will say my inbox is blowing up from every tool under the sun that I use, letting me know that I can now use Opus 4.8 in them, just like on Ritual. And I will say, I think that for our track record, this is now the second consecutive Claude model drop that's happened literally as we've stepped in here to talk about the weekly news. It's almost like they know when our segment is, then they're just trying to make sure we're slightly outdated or something. So Anthropic, if you're listening to this, just give us a heads up if you're about to drop it, because we're seeing to be aligned on our timing.

**Ben Lloyd-Pearson** (0:47)
Yeah, it'll be interesting to try it out. I mean, it doesn't seem like they're really making a whole lot of groundbreaking claims with this one, compared to like Mythos, for example.

**Andrew Ziegler** (0:55)
Yeah, it's maybe a more routine upgrade of a model that we've all grown pretty adapted to. And I will say, there's flocks of people running towards these models right now, but there's also a lot of folks that are going against the grain, maybe going the other direction, and really revisiting local models and local infrastructure, and maybe not making the same bets on, oh, these anthropic models, oh, these open AI models, these frontier models are going to get better and better. I should put all of my work and time on them. I think what we're seeing is the reality of using them at scale, and the cost is like a pot of boiling water, and we're all frogs in it right now.
What do you think?

**Ben Lloyd-Pearson** (1:33)
Well, we're going to get in to that. It's definitely one of the topics we're going to cover today. So welcome to the Friday Deploy brought to you by LinearB. I'm your host, Ben Lloyd-Pearson.

**Andrew Ziegler** (1:42)
And I'm your host, Andrew Ziegler.

**Ben Lloyd-Pearson** (1:44)
And this week, we are covering moving your AI models local, AI data centers and how they're changing the needs of infrastructure. AI as a workflow crutch, the eternal slop timber, and we're all tired of AI conversations, even the bots potentially. So, Andrew, let's get into it. You mentioned moving things local.
What's this article that we have about outsourcing and using local AI to take a more economical approach to AI usage?

**Andrew Ziegler** (2:15)
Yes. If you've been using Inference, and if you're listening to this podcast, you most likely are, you've noticed that your bill has been starting to climb. The API consumption costs at major frontier model providers is slowly escalating, as I mentioned at the top of our show.
And this is pushing a lot of teams and a lot of leaders that have now invested and built AI infrastructure to reconsider what models they're betting on.
A lot of folks are actually running the opposite direction from these new model releases and fine-tuning local models like Quen 3 and a whole bunch that are Apache 2 at this point. You fine-tune them on your own local data, you host them on your own infrastructure, run them on your own GPUs, eliminates a lot of the uncertainty about budgeting token costs for the future. And what's really interesting here is that our token consumption continues to climb. And if the token costs are going to start slowly increasing at the same time, you're going to get this compounding increase in price that's going to sneak up on a lot of people. So that's why we're seeing these large organizations get really strategic about it early. We talked about Shopify fine-tuning an agent that then allowed them to build out a whole internal multi-agent infrastructure they couldn't even afford before on the inference when they were using it off the shelf from a foundation model. So I just want to remind folks that the technology is out there to own this yourself, to figure out what it would mean to have your own specialized intelligence, your own specialized harness. You don't have to rely on just the top model providers. And we're going to be in a world soon where model routing and choosing when to use a quote dumb route, a dumb model or quote a smart model is going to become a more important choice than ever has been before. So those are some of the moves that I'm seeing contrary to a lot of these new model drops happening like Google last week as well.

**Ben Lloyd-Pearson** (4:14)

31 more minutes of transcript below

Thousands of transcripts fetched by people building searchable podcast archives

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/YOUR_EPISODE_ID