A deep dive into Nvidia's Vera CPU and the Olympus cores that power it

sbulaev1 pts0 comments

Nvidia's Vera CPU and the Olympus cores that power it: Deep dive

Jump to main content

Search

REG AD

SYSTEMS

A deep dive into Nvidia's Vera CPU and the Olympus cores that power it

88 custom cores, 176 funky threads, 1.5 TB of laptop RAM, and 1.8 TB/s of NVLink connectivity — this isn't your typical datacenter chip

Tobias Mann

Tobias<br>Mann

SYSTEMS EDITOR

Published<br>sat 1 Aug 2026 // 10:02 UTC

DEEP DIVE For the first time, Nvidia has directly challenged Intel and AMD's CPU dominance. With the launch of Vera, the AI arms dealer aims to flog its standalone CPUs to as many hyperscalers and other cloud providers as it can.<br>Alibaba, ByteDance, Meta, Oracle, CoreWeave, Lambda, Nebius, and NScale have already signed up to deploy the chips in their respective clouds.<br>The follow-on to Nvidia's Grace CPU promises 88 custom Armv9.2 cores, 176 threads, support for up to 1.5 TB of LPDDR5X memory, and, critically, availability as a standalone platform independent of Nvidia's GPUs.

REG AD

But beyond that, and a mountain of marketing about how it'll be the best CPU for everything AI, Vera's inner workings have largely remained a mystery until recently.

REG AD

That changed late last month, when Nvidia released a whitepaper spilling the beans on its first fully-custom CPU, which is far weirder than anyone could have anticipated.<br>Ostensibly, Nvidia is aiming Vera at two key workloads: the first and least surprising is as the AI head node responsible for managing the GPUs in its upcoming Vera Rubin systems. The second, and more contentious, is as a host for AI agents, which, unlike the large language models (LLMs) that power them, don't actually run on GPUs.<br>From what we gather, much of Vera's core architecture is predicated on quashing pipeline and execution bottlenecks in order to make it more effective in these roles. But before we dive into Nvidia's Olympus core, let's revisit the chip itself.<br>Monolithic compute, multi-die memory and I/O

Here's a quick refresher on Vera's construction.<br>Image credit Nvidia

Peel back Vera's rather substantial heat spreader and you'll find an assortment of chiplets responsible for I/O, memory, and compute. However, this isn't another rehash of the chiplet architecture popularized by AMD.<br>Unlike x86 processor makers, which spread dozens of cores across multiple dies, Vera's compute die is monolithic. All 88 cores are housed in one big chunk of silicon that from what we understand is fabbed on TSMC's 3nm process. Nvidia argues its monolithic compute architecture has benefits for both core-to-core bandwidth and latency compared to competing designs.<br>Surrounding the compute die is an assortment of chiplets, including eight LPDDR5x controllers, and what appears to be two distinct I/O dies, one responsible for PCIe 6.4 and CXL 3.1 connectivity and another dedicated to the chip's NVLink Chip-to-Chip interface.<br>In terms of package design, Vera is somewhat reminiscent of Amazon's Graviton 4 CPUs, which also combine a monolithic compute die while disaggregating the I/O and memory functionality to dedicated silicon.

REG AD

The Vera CPU Superchip

Nvidia’s Vera CPU superchip packs 176 cores and 3 TB of LPDDR5x memory.<br>Image credit Nvidia

Like most modern datacenter CPUs, Vera supports both single- and dual-socket configurations, the latter of which Nvidia calls the Vera CPU Superchip.<br>Unveiled at GTC in March, the Superchip features two Vera CPUs connected over NVLink-C2C at a blisteringly fast 1.8 TB/s of bidirectional bandwidth, for a total of 176 cores and 352 threads.<br>The two chips are fed by 16 SOCAMM2 LPDDR5x memory modules, which deliver 2.4 TB/s of aggregate memory bandwidth (1.2 TB/s each) — roughly twice the bandwidth of AMD's Turin Epycs, which were launched in 2024.<br>Nvidia's agentic AI reference designs call for cramming as many as 128 of these superchips (256 CPUs) totaling 22,528 cores and 384 TB of memory into a single liquid-cooled rack.<br>Summiting Olympus<br>Previously, Nvidia had relied on off-the-shelf CPU cores from Arm. For example, depending on the iteration, Nvidia's Grace CPUs used Arm's Neoverse V2 (GB200/300), Neoverse V3 (AGX Thor), or Cortex X925 and A725 cores (GB10), depending on which was most convenient.<br>With Vera, Nvidia has moved on to designing its own ARMv9.2-compatible core, called Olympus. In fact, it's up for debate how custom the chip really is. We've been told that the chip builds heavily on existing Arm IP, and looking at Olympus' architectural block diagram, we can see why.

REG AD

Here's a block diagram of Nvidia's Olympus core architecture.<br>Image credit Nvidia

Olympus features a 10-wide decoder and dispatch, eight integer arithmetic logic units (ALUs), six vector/FP pipelines (SVE 128), four load units, and 2x store units. So it's got a fat front and back end compared to the Zen 5 cores found in AMD's Turin Epycs or the Redwood Cove cores used by Intel's Granite Rapids Xeons.<br>But compared to off-the-shelf Arm cores, Olympus' front and back end...

nvidia vera cores olympus memory chip

Related Articles