Tech Bytes: Bring AI Inferencing on Prem with VMware AI Factory (Sponsored)
The Fat Pipe - All Packet Pushers Pods
September 28, 2026
The early exuberance of AI adoption has run up against hard-nosed business requirements around managing costs, protecting sensitive data, and matching the right model with the right job at the right price.
Speakers Packet Pushers, Sabina Anya
TopicsTechnology
Packet Pushers (0:04)
Welcome to TechBytes. So the early exuberance of AI adoption is now running up against hard-nosed business requirements around things like managing costs, protecting sensitive data, and matching the right model with the right job at the right price. On today's show, sponsor Broadcom is here to make the case that private cloud is the place for AI inferencing workloads. We're gonna talk about the new VMware AI factory, which combines software and hardware to accelerate deployment and operations. There's also curated models to make it easy to deliver inferencing as a service to our internal customers. And we'll talk about some more stuff too. My guest is Sabina Anya. She is Chief Technologist and Executive Advisor for VMware Cloud Foundation.
Sabina, welcome. And VMware recently announced new options for organizations that want to run AI inferencing workloads in private clouds. What's driving this, I guess, sort of repatriation back to private cloud?
Sabina Anya (0:51)
Hello there. I think one of the things we want to touch on first is some of the business drivers. We had something like a 97 percent of IT leaders saying that public cloud has no public cloud spend has become wasted. All right. So you have a lot of budget that you're investing into this. If all of these challenges with like the AI act, are my workloads safe in the cloud? Can somebody tamper with them and so on? That basically triggered a certain amount of, can we do this on premises? Does it make the most amount of sense? Is it the most predictable maybe on private cloud? So based on that, we get into this space where building on premises with VMware technology that you know and love becomes a lot easier and a lot more obvious than it used to be.
Packet Pushers (1:40)
So one of the things we hear about AI cost is around tokens. Cost of tokens is a cost driver.
I'm curious how AI consumption costs compare, like if I'm using an external service versus the capex and opex costs associated with me having to deploy and operate a private cloud.
Sabina Anya (1:58)
You know, I love that statement because our chief product officer, Paul Turner, went on stage just last week and said basically, no more tokens, right? Because every company, it's that whole no more tokens, literally bar through it, don't think about it. And not in the sense that tokenomics is not an important tool.
It's still how we measure performance and all of that, right? But the problem is that if you think of how every company uses AI, you basically have this full, you have to use it because it makes you more efficient, but at the same time, you're constantly gated. You can only use this much, you can only do this much, you can only, so you have all of these limitations of cost per millions of tokens. Which model can I use? How much can I live there? And you have all of these overly inflated costs at some point that come from, did I need that? Did I set it up correctly? Whereas on private cloud, you set it up with what you have, and you get into this point where, I would say you consume the most of everything, especially when you do the architecture properly and correctly. So there's this whole discussion on how do you invest in private cloud infrastructure in a way that allows you to make the most use of what you have. It's a very complex way of saying maximum ROI at the end of the day, but you told me it can get nerdy, so I'm not going to go there.
Packet Pushers (3:16)
Please do. Okay, so you're saying despite capex cost, opex cost, do you think there's a financial argument can be made that private cloud can be more efficient, maybe even more economical?
Sabina Anya (3:31)
I think private cloud can be more bespoke to your needs. Now, that obviously triggers a different kind of challenge as well, which is setting it up according to your needs. You can't just deploy AI on top of it and hope for the best. Some of the recent statistics show that GPU utilization is still around 30 percent, especially when used in bare metal environments. That's very far away from what I'm saying right now. So you need to get into a space where, what do I actually need? What are the workloads that I'm going to deploy? How is this AI infrastructure going to make use of it? And how do I get the maximum insight into what I have deployed and how it uses everything? So I can go say, okay, I need better time-sharing and time-slicing of CPUs. I need better usage of my network infrastructure, for example, and instead of having ports that have no forwarding, how do I push more traffic out? All of these things become key to designing a data center for the AI era.
14 more minutes of transcript below
Thousands of transcripts fetched by people building searchable podcast archives
Fetch the whole transcript
The demo key returns a sample episode in full, no card needed:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090Markdown with the speakers named, for your notes, your knowledge base, or anything that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000792072120