AMD attacks the rack with Helios systems that rival Nvidia's

galaxyLogic1 pts0 comments

AMD attacks the rack with Helios systems that rival Nvidia's

Jump to main content

Search

REG AD

Systems

AMD attacks the rack with Helios systems that rival Nvidia's

Spec for spec, the House of Zen's first rack-scale AI compute platform is bigger and faster than Nvidia's Vera Rubin by nearly every metric, but that's only on paper

Tobias Mann

Tobias<br>Mann

SYSTEMS EDITOR

Published<br>thu 23 Jul 2026 // 18:26 UTC

Could Nvidia's days of datacenter dominance be threatened? With the launch of Helios, AMD’s first true rack-scale AI platform, the House of Zen is done playing catch up.<br>The company claims the rack system is powered by the fastest AI accelerators on the market. And, at least on paper, the 72-GPU system is not only bigger and faster by nearly every metric than Nvidia’s existing Blackwell-based rack systems but than Vera Rubin as well, and that includes the rack.<br>Measuring in at 1.2 meters wide and 44OUs high, the OCP Open Rack Wide form factor is nearly twice the size of Nvidia's NVL72, and AMD has clearly put the extra space to good use.

REG AD

Compared to Vera Rubin, Helios boasts 50 percent more HBM4 and scale out bandwidth and between 15 and 25 percent higher performance for AI training.

REG AD

Vera Rubin’s adaptive compression tech will supposedly give it a 25 percent lead over Helios at FP4, but, as we understand it, that’s only for inference workloads. For applications that can’t, Helios offers 15 percent higher peak FP4 FLOPS.<br>To be clear, it’s not the first time we’ve seen AMD pull ahead on memory or FLOPS. The difference is those products usually launched a year after Nvidia’s equivalent part. Helios launches right in time with Nvidia’s Vera Rubin platform.<br>AMD estimates Helios' higher peak performance will give it a 30 percent performance per dollar lead over the competition.<br>The belly of the beast<br>Helios' performance gains are rooted in an all-new GPU built on AMD’s 5th-gen CDNA compute architecture.<br>Much like the rack it powers, the Instinct MI455X is massive, though the chip is underselling it a bit. Just like the MI300 series, AMD’s latest datacenter GPU is a silicon sandwich that stitches together I/O, compute, and memory in a single package.

Like previous Instinct GPUs, the MI455X isn't one chip, but nearly two dozen compute, I/O, and memory chiplets stitched together with TSMC's advanced packaging tech.<br>Image credit AMD

Including memory, MI455X features 24 chiplets using a combination of 2.5D and 3D packaging technologies. The chip’s eight compute dies are fabbed on TSMC’s bleeding edge 2nm process tech, which are stacked atop a pair of 3nm fabric and cache dies (FCDs). The FCDs are an interesting twist on the formula. They function as a cache heavy interposer with 96 MB of L2 cache each and the memory controllers for the chip’s 12, 36 GB HBM4 stacks.<br>Unlike past Instinct accelerators in which the I/O die was located under the compute, the MI455X breaks these out into two new dies — also fabbed on TSMC’s 3 nm — which are responsible for chip-to-chip communication.

REG AD

AMD's chiplet architecture means it can function as one big GPU, two smaller ones, or up to eight virtualized accelerators.<br>Image credit AMD

One benefit to this architecture is that the chip can be made to function as one big GPU or two smaller ones depending on which NUMA configuration you opt for. The chip also supports spatial partitioning into up to eight virtual GPUs.<br>Under the hood, AMD’s CDNA architecture brings some notable improvements over the last generation. Compared to last year’s MI355X, the MI455X promises as much as 4x higher floating point performance for AI workloads.

In just a year, AMD has managed to deliver a 4x uplift in floating point performance over MI355X.<br>Image credit AMD

Most notably, the MI455X forgoes FP64 entirely in order to dedicate as much die area to AI-centric datatypes like MXFP4 and MXFP8 as possible. This generation also adds support for 16 and 32 block scale data types. For those needing FP64 compute, that functionality will be served by a different HPC centric SKU.<br>This isn’t the only physical change. As we mentioned earlier, for this generation, AMD has opted for a larger shared L2 cache and ditched its last level “Infinity” cache altogether.<br>The benefit, AMD fellow Alan Smith says, is higher bandwidth and a simplified data path compared to last gen. “The bandwidth delivered for one of these L2 caches in MI455X is 1.5x the aggregate bandwidth of the Infinity cache on MI355X."<br>Along with prioritizing low precision compute, AMD configured the chip’s execution engines to boost IPC and implemented a new direct memory access (DMA) engine to minimize data movement.<br>Racking it up<br>While bigger than Nvidia’s NVL72, AMD’s Helios rack architecture is remarkably similar.

REG AD

Here's a quick rundown of the Helios compute blade, which crams four MI455X GPUs, 12 Vulcano smartNICs, a Salina DPU, and a 96-core Venice Epyc CPU into a 1OU form factor.<br>Image credit...

helios rack nvidia compute chip mi455x

Related Articles