OpenAI Pauses Model After Sandbox Escape artwork

OpenAI Pauses Model After Sandbox Escape

The Daily AI Show

July 21, 2026

The episode opened with Kimi K3, Qwen 3, and the practical limits of open weight frontier models. The hosts discussed why these Chinese models may be cheaper to use through hosted inference, but still require massive data center resources to run directly.
Speakers: Beth Lyons, Andy Halliday, Gareth
**Beth Lyons** (0:00)
Hey, good morning, everybody. It is Tuesday, July 21 It is episode maybe 777, something like that.
I feel like we're moving toward the seven eights. And I'm Beth Lyons with me in the studio today, is Andy Halliday. We'll see if anyone else pops in the back door. How you doing, Andy?

**Andy Halliday** (0:24)
I'm well, thank you.

**Beth Lyons** (0:25)
Awesome.
You are watching and listening to The Daily AI Show, which I need to remember to say when in fact, that is true and we are live. Andy, some interesting things. I didn't see a big model drop yesterday, but I did see a lot of conversation that is getting to interesting places. What did you bring today?

**Andy Halliday** (0:51)
Well, related to model drops, there is a lot of discussion on this show and elsewhere about Kimi K3 and also Qwen 3.8, those Chinese models with much lower operating costs for enterprises or other.
But one of the obstacles to that is these are very large models like over two and a half trillion parameters. That means that, for example, if you have an innate substack, he made the point that it's open weight, yeah, sure, but doesn't mean that you can actually operate it. This is really only operable as an open source model by somebody who has a very large data facility. Specifically, it takes about 64 high-end GPU systems on a chip. So like a GB300 from Nvidia, you have to buy 64 of those.
That's not $1,000, that's not $10,000, not $100,000 worth of chips. That's way more than that in order to run such a thing. So these open models are available to data centers. So if you operate a data center, you can run one of these models. But anyway, I digress because what I really wanted to say is that there are credible reports out that Microsoft is aggressively exploring using Kimi K3 as one of the network of models, including their MAI, Microsoft AI models internally to replace OpenAI and Anthropic, which they consider to be inherently expensive. Why Mustafa Suyman, who is the head of Microsoft AI has publicly said that Anthropic is extremely expensive and that Microsoft's goal is to, and I quote, reduce and ultimately eliminate that cost by using its own MAI models and tools that are available to them like Kimi. Now, Microsoft obviously has major data center capacity, so it can run Kimi K3 for its own customers and for its own internal use. It could also run a host of other models. So is this the end of the primacy of OpenAI and Anthropic, or is it just there are companies scrambling to solve the compute shortage that exists out there right now, and ultimately prices will stabilize, but right now we're sitting under some umbrella pricing, actually, for some of the frontier model developers, where their advanced capabilities, this capability overhang that we sometimes talk about, are very appealing to people doing development using these AIs and or architecting agentic systems that can use the main frontier intelligence as the orchestrator for a host of less expensive agents, and really not incur huge inference costs, as if everybody in the company is using Fable 5 or Sol 5.6 as a chatbot.

**Beth Lyons** (4:24)
Right. And I actually want to make a little distinction in what you said, because actually anyone could download Kimi K3, but you need something like Microsoft's resources to be able to run it. Because there is a little bit of a conversation around the Trump administration was making noise about, hey, maybe we don't allow frontier Chinese models that are open source, open weight to be available. But of course, the weights aren't open yet, but it is available now.
Actually, Matt Wolf, I think, was the person that I was reading, was like, okay, so does this mean download it now, even if we can't run it? What is the recommendation here?

**Andy Halliday** (5:20)
If you download the weights, how many gigabytes is that?
Let me put it this way. I downloaded Gemma 4 billion, like 4 billion parameters. I think it was, I don't know, 175 gigabytes or something for that. I can't do the math in my head, but you're going to have to have a massive machine just to download this thing.

**Beth Lyons** (5:50)
Well, I think yes, but compared to cores and RAM and superchips, the ability to put some gigabytes together to be able to host something is significantly- Yeah, what are you going to do with it?

**Andy Halliday** (6:04)
It's like, soon be out of date. You're going to continuously update that in some ways? I don't know.

**Beth Lyons** (6:12)

41 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000777780034