**Skyler Monroe** (0:00)
Welcome back to the AI Hardware Show. I'm your host Skyler Monroe, and I'm here to break down everything happening in the wild world of AI chips, from GPUs to custom silicon. Before we dive in, huge thanks to our sponsors. LimitLess AI, helping clients seamlessly integrate AI into their workflows, and a Go Consulting, that's a Go, your go-to for Silicon Development Consulting. Today, we're covering some massive moves in the chip world. Nvidia's surprising deal with Groq. Blackwell Jitka is hitting Enterprise Data Centers, some killer benchmarks comparing the latest Silicon, and Meta's custom processor reveal. Let's jump right in. All right, let's start with what might be the most surprising news of the week. Nvidia apparently making a deal with Groq. Now, if you're thinking, wait, isn't Groq supposed to be Nvidia's competitor? You're absolutely right to be confused. Groq has been positioning itself as the anti-GPU company, focusing specifically on low-latency inference with their language processing units, or LPIs. These aren't your traditional GPUs. They're purpose-built for running AI models as fast as possible once they're already trained. Think of it like this. If training an AI model is like building a race car, inference is actually driving that car in the race. Nvidia has been dominating both the car building and the racing part, but Groq said, hey, we can make a better race car driver.
Their LPUs are designed with a completely different architecture that prioritizes speed over everything else. While Nvidia's HA100S are like Swiss Army knives, incredibly powerful and versatile. Groq's chips are more like Formula One cars, built for one thing and one thing only, going really, really fast. What's fascinating about this potential deal is what it signals about the market. Nvidia isn't just the big bad monopolist trying to crush everyone. They're actually validating that there's room for specialized players. It's like McDonald's deciding to partner with a high-end burger joint instead of trying to compete directly. The inference market is massive and growing. And even Nvidia recognizes they can't be everything to everyone. This move validates the entire non-GPU AI chip startup ecosystem. Companies like Cerebras, Sambinova and Graphcore have been arguing for years that specialized silicon can outperform general purpose GPUs for specific tasks. And now even Nvidia seems to agree. The fallout from this Nvidia-Groq deal is sending shockwaves through the entire AI chip startup landscape. And honestly, it's about time the industry had this wake up call. For the past few years, we've seen hundreds of AI chip startups all claiming they're going to dethrone Nvidia. But most of them have been fighting the wrong battle. They've been trying to build better training chips, which is like trying to outmuscle the rock at his own game. But what this deal shows is that Nvidia gets it. The future isn't just about who can build the biggest, most powerful chip. It's about building the right chip for the right job. Training and inference are fundamentally different workloads. Training is like construction. You need raw power, lots of memory, and you're okay if it takes a while.
Inference is like a restaurant kitchen during dinner rush. You need speed, efficiency, and consistency above all else. The validation here isn't just for Groq. It's for the entire concept of workload-specific silicon. We're seeing this play out across the industry. Companies like Cerebras focusing on massive scales training, Samba Nova targeting enterprise inference, and now Groq potentially partnering with Nvidia for ultra-low latency applications. What's really interesting is how this changes the competitive landscape. Instead of every startup trying to be the Nvidia killer, we're moving toward a more mature ecosystem where different companies excel at different things. It's like how the car industry evolved. You don't see Ferrari trying to make minivans or Toyota trying to make supercars. Each company focuses on what they do best. This deal also sends a clear message to investors. Stop looking for the one company that's going to replace Nvidia everywhere. And start looking for companies that can beat Nvidia at specific high-value use cases. Now, let's talk about Nvidia bringing their Blackwell GPUs to enterprise data centers. Because this is where things get really interesting from an infrastructure perspective. Blackwell isn't just a new chip. It's a complete rethinking of how we build AI compute infrastructure. The GAB200 Grace Blackwell Superchip combines two B200 GPUs with Nvidia's Grace Kipu, and the numbers are just insane. We're talking about chips that can deliver up to 20 petaflops of AI compute performance. To put that in perspective, that's more compute power than entire supercomputers had just a few years ago. And now it's in a single server node. But here's what's really clever about Blackwell.
Nvidia isn't just making faster chips. They're completely reimagining how these chips work together. The NvLink connections between chips can transfer data at 1800 gigabytes per second. That's like having a fire hose of data flowing between chips. Traditional server architectures are like city streets with traffic lights. Data has to stop and wait all the time. Blackwell is like building super highways with no traffic lights where data can flow at maximum speed all the time. What's particularly smart about the enterprise rollout is how Nvidia is packaging this. They're not just selling chips. They're selling complete rack scale solutions. The NvL72 systems pack 72 Blackwell GPOs into a single rack. And these things are basically super computers in a box. Enterprise customers don't have to figure out cooling, networking or system integration. Nvidia has done all that work. It's like buying a fully loaded sports car instead of buying an engine and trying to build the rest yourself. The liquid cooling requirements alone would be a nightmare for most enterprises to figure out on their own. These systems generate so much heat, the air cooling just doesn't cut it anymore. But the real game changer here is what this means for AI deployment timelines. Companies that used to take months to deploy AI infrastructure can now have pediscale compute up and running in weeks. Let's dive into some hardcore benchmarking with the InferenceX V2 results. Because these numbers tell the real story of where AI hardware performance is heading. We're looking at Nvidia's new GB300 Avial72 system going head to head with ARD and these ME355X and comparing against the current Hopper generation. What's immediately striking is how much the performance landscape has shifted with these new architectures. The GB300 system isn't just incrementally better. It's showing performance jumps that are frankly mind-blowing for certain workloads. For large mixture of experts models, we're seeing throughput improvements of 2 to 3x over previous generation hardware. Think of mixture of experts models like having a panel of specialists instead of one generalist. Each expert handles what they're best at, but you need massive parallel processing power to coordinate them all effectively. What's really interesting in these benchmarks is how different architectures excel at different things. AMS My355X is showing surprisingly competitive performance in certain inference scenarios, especially when you factor in the cost per token generated. It's like comparing different race cars on different tracks. The Formula 1 car dominates on the speed track, but the rally car might win on the rough terrain. The disaggregated serving results are particularly fascinating because they show how modern AI workloads are becoming more distributed. Instead of running everything on one massive chip, we're seeing better performance by spreading the workload across multiple smaller specialized processing units. Sclong and Vellum optimizations are showing dramatic improvements in serving efficiency. These aren't just minor tweaks. We're talking about 40 to 60% improvements in tokens per second for real-world serving scenarios. The wide expert parallelism results really highlight how important memory bandwidth has become. It's not just about compute anymore. It's about feeding that compute fast enough to keep it busy. And finally, let's talk about Meta's reveal of their custom AI processors for data center operations. Because this is a perfect example of how the biggest tech companies are moving beyond just buying chips off the shelf. Meta isn't just designing one custom chip. They're building an entire portfolio of specialized processors optimized for their specific workloads. This is like a restaurant chain deciding to build their own custom ovens instead of buying commercial ones. Because their cooking needs are so specific that generic equipment doesn't cut it anymore. What's particularly interesting about Meta's approach is how they're thinking about the entire stack. They're not just designing inference chips or training chips. They're building processors optimized for recommendation engines, content moderation, video processing, and even Reality Lab's applications. Each of these workloads has completely different computational requirements. Recommendation systems need to process millions of user interactions in real time with extremely low latency. Content moderation needs to analyze text, images and video simultaneously. VER and AR applications need specialized computer vision and spatial computing capabilities. The smart move here is that Meta isn't trying to build one chip that does everything. They're building a family of chips that work together. It's like having a toolbox where each tool is perfect for a specific job, rather than trying to do everything with a single multi-tool. This also gives Meta incredible cost advantages. When you're operating at the scale of billions of users, even small efficiency improvements translate to millions of dollars in savings. Plus, they're not paying the Nvidia tax, that premium you pay for buying the best GPU on the market. But the real strategic advantage is control over their own destiny. Instead of waiting for chip vendors to build what they need, Meta can optimize their hardware exactly for their software and their use cases. This is the same playbook that made Google's TPUs so successful for their search and advertising workloads. That's a wrap on today's episode of The AI Hardware Show. We covered some massive shifts in the chip landscape this week, from Nvidia's surprising partnerships to Enterprise Blackwell deployments, and Meta's custom silicon strategy. The AI Hardware Show brought to you by LimitLess AI and a Go Consulting. Make sure to subscribe wherever you get your podcasts. And if you found this useful, share it with fellow hardware enthusiasts. All the links and detailed specs we discussed today are in the show notes. Until next time, keep those chips running cool and those benchmarks high.
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000754787984