Topics: Technology, News, Tech News
**Doug Black** (0:05)
Welcome to HPC News Bytes, a weekly show about important news in the world of supercomputing, AI, quantum computing, and other advanced technologies. Hi everyone, welcome to HPC News Bytes. I'm Doug Black, and with me of course is Shahin Khan.
There's a rising tide of GPUs on the AI processing market, many of them from some of the richest companies in the world, but the degree to which they pose a threat to NVIDIA's GPU market dominance is very much in question. There was news last week that Meta plans to have its new AI chip in production in the September timeframe, according to a Reuters story. The report is based on an internal Meta memo. The new chip is called Iris, announced four months ago, and it's the first in a series of Meta training and inference accelerator processors to be rolled out over the next year and a half, and they will target generative AI inferencing workloads. Meta's Iris chip is one piece of a larger move of Meta and other hyperscalers toward owning and optimizing more of their own AI infrastructure stack.
**Shahin Khan** (1:16)
That really is the big trend and the big signal here. Broad brush, if you have enough volume, you can build your own chips and systems. Hyperscalers can build their own chips and systems and software, and they increasingly do, while continuing to buy from merchant vendors what they must.
Chip vendors, likewise, are becoming system and rack scale vendors, and invest in cloud providers who are strongly aligned with them.
Meta is working with Broadcom on custom silicon and TSMC on fabrication and packaging, while continuing to buy heavily from Nvidia, AMD, and Intel for merchant compute.
The supporting supply chain is just as important. Samsung, SK Hynix, Micron, and SanDisk for memory and storage. Sumitomo Electric for optical interconnect components. WeWin, Quanta, and Foxconn for servers, motherboards, and racks. The layers point to another market signal here. Hyperscalers building their own hardware and the growing complexity of high-end chips are coming together to change the supply chain for chips. Broadcom, Marvell, and increasingly MediaTek are moving up the stack from implementation partners towards full-service custom chip prime contractors. Hyperscalers increasingly specify the workloads, economics, and system requirements, while specialist partners take on more of the architecture, design, packaging, and production. That could become a major new layer in AI infrastructure.
**Doug Black** (2:50)
The Oak Ridge Leadership Computing Facility announced last week that they're deploying a 20-qubit system from IQM Quantum Computer.
Oak Ridge said the IQM Radiance System, which is named Pathfinder, will play a key role in the lab's efforts to integrate quantum computing technology with their classical HPC resources. IQM was founded in 2018 and is based in Finland. Their Radiance System is based on superconducting technology, which means its qubits must be cool to nearly absolute zero.
**Shahin Khan** (3:25)
The Oak Ridge installation reinforces a broader shift already on their way in the US. Quantum computing is moving from experiment and engineering into procurement and operations. The 20-qubit system from IQM is modest. They have bigger systems that they sell to. But the strategic value is in integrating quantum hardware with HPC workflows, schedulers and applications while developing in-house expertise. The larger market signal is also important. We can now identify more than 30 quantum computing vendors across 8 hardware modalities that have collectively built or deployed well over 100 systems, including internal research platforms, prototypes, cloud systems and customer installed machines.
IQM for its part has said that it has sold 23 systems and installed 15
**Doug Black** (4:18)
HPC luminary Satoshi Matsuoka of Japan's Riken Center for Computational Science has issued three papers of late that, given his stature in the supercomputing industry and the content of his articles, has stirred the pot in the HPC community. Matsuoka, of course, is a leader in Japanese supercomputing system strategy and design, including the country's upcoming Fugaku NEXT leadership system, scheduled for 2029 or 2030 His three papers, all available on the archive site, take on the topics of floating point 64 processing as the HPC Holy Grail. Another argues that FP8 is all you need. The third looks at the scarcity of memory chips at open AI models, and at the restructuring of the AI industry through the rest of the decade.
**Shahin Khan** (5:11)
We need to cover these topics in more detail, but I will share my high level take from these papers. On numerical precision, if you assume that most scientific computation can be formulated as matrix and tensor algebra, a reasonable assumption, then algorithmic emulation can do the job faster. If the fraction of computation that is not matrix algebra is small enough, then it can be done in slower ways and you still come out ahead. On the memory issues, I see the paper making the following salient points. A trend towards memory bandwidth versus compute. Memory bandwidth is what people are in fact buying, the paper says. Older infrastructure can continue to work well and could dampen the need to upgrade and especially in situations where the infrastructure that runs a service is not visible to the customer.
3 more minutes of transcript below
Thousands of transcripts fetched by people building searchable podcast archives
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/YOUR_EPISODE_ID