Nvidia’s Vera vs Intel and AMD: 88 Olympus cores, 1.8TB/s NVLink, and standalone CPUs
Nvidia’s first fully-custom Armv9.2 design targets agent-host workloads and hyperscale head-nodes, without relying on its GPUs.

Nvidia is pushing Vera, a standalone CPU platform with 88 custom Armv9.2-compatible Olympus cores, after previously using off-the-shelf Arm cores in Grace. The bet is that hyperscalers will deploy Vera as both a GPU-managing head node and as AI agent infrastructure.
Nvidia has decided to pick a fight the CPU market usually avoids: challenging Intel and AMD’s dominance with a standalone chip, not a GPU-adjacent curiosity. Vera, Nvidia’s first fully-custom CPU, aims straight at hyperscalers and cloud providers with a system built around 88 custom Armv9.2-compatible Olympus cores, 1.8 TB/s of NVLink connectivity in its “Superchip” configuration, and a design that is available as a CPU platform independent of Nvidia’s GPUs.
And it is already lining up customers. Alibaba, ByteDance, Meta, Oracle, CoreWeave, Lambda, Nebius, and NScale have signed up to deploy the chips in their respective clouds. The early positioning matters because it clarifies what Nvidia thinks the CPU’s job is inside an AI stack: not just compute for LLMs, but orchestration for where GPUs bottleneck and infrastructure for “AI agents,” which do not run on GPUs the way large language models do.
Vera’s outward pitch is big, but the real story is how Nvidia tries to eliminate the kinds of stalls that silently drain performance. The company says Vera is designed around quashing pipeline and execution bottlenecks to improve its effectiveness in those roles. It targets two workloads. First, the less controversial one: Vera as the AI head node that manages GPUs in Nvidia’s upcoming Vera Rubin systems. Second, the contentious one: Vera as a host for AI agents, where code is often branch-heavy and execution can become a game of waiting for the next instruction.
Under the heat spreader, Vera is a “superchip” style system that is still recognizably modern datacenter silicon, but with a twist: its compute die is monolithic. Nvidia argues that having all 88 cores on one chunk of silicon improves core-to-core bandwidth and latency versus competing multi-die compute approaches. Nvidia says the compute die is fabbed on TSMC’s 3nm process. Around that monolithic compute die are dedicated chiplets for I/O and memory, including eight LPDDR5x controllers, plus what appears to be two distinct I/O dies: one for PCIe 6.4 and CXL 3.1 connectivity, and another dedicated to NVLink Chip-to-Chip interface.
That memory bandwidth story is part of why Vera sounds built for AI datacenters instead of traditional enterprise workloads. In its single- and dual-socket modes, Vera supports the kind of aggressive bandwidth modern systems require. In the dual-socket configuration Nvidia calls the Vera CPU Superchip, the Superchip pairs two Vera CPUs connected over NVLink-C2C at 1.8 TB/s bidirectional bandwidth. That gives a total of 176 cores and 352 threads. The two chips are fed by 16 SOCAMM2 LPDDR5x memory modules, delivering 2.4 TB/s aggregate memory bandwidth, or 1.2 TB/s each. Nvidia also frames this against competitors, saying that 2.4 TB/s is roughly twice the bandwidth of AMD’s Turin Epycs launched in 2024.
Agent-hosting and head-node orchestration are not just about raw bandwidth. They are about CPU execution behavior, especially when software is pointer-heavy, graph-like, or generated on the fly. That is where Nvidia’s custom core, Olympus, comes in. Previously, Nvidia relied on Arm’s existing cores for Grace (for example, Arm’s Neoverse V2, Neoverse V3, or Cortex X925 and A725 depending on the iteration). With Vera, Nvidia shifted to designing its own ARMv9.2-compatible core called Olympus.
The interesting part is how “custom” Olympus is claimed to be. The block diagram shows Olympus has a 10-wide decoder and dispatch, eight integer ALUs, six vector/FP pipelines (SVE 128), four load units, and 2x store units. In other words, it has a fat front and back end compared to Zen 5-style designs from AMD’s Turin Epycs or Intel’s Redwood Cove cores in Granite Rapids Xeons, but it is not wildly unrecognizable. Nvidia’s argument for differentiation centers on features aimed at performance under branch-heavy and stall-prone workloads. The company claims Olympus includes an entirely custom neural branch predictor capable of exploring two branches simultaneously, which it says reduces mispredict likelihood and boosts performance.
Nvidia also says it addressed pipeline stalls through mid-core changes. Memory renaming is used to speed up store-to-load dependency chains, letting dependent instructions execute before a load completes if the relationship to the data can be inferred. Nvidia calls out pointer-heavy software, graph traversal, runtime frameworks, and complex object-oriented workloads. It also implemented a value prediction scheme that identifies stable dependency chains and predicts future values before they are produced, allowing dependent instructions to execute speculatively with correctness verified later. Nvidia positions that as helpful for repetitive software patterns and sequential data processing.
On the back end, Olympus keeps the familiar datacenter mix of arithmetic and vector capability, but Nvidia says it implements some things differently. It has eight ALUs split into simple integer units for basic mathematics and two complex ALUs for multiplication and division, CRC, and shift intensive workloads. There are four dedicated branch units tied to the branch predictor to reduce stalls from control flow decisions. For floating point, the core includes six 128-bit Arm SVE2 extensions supporting FP8, plus a pair of vector units dedicated to speeding up crypto operations. Memory-wise, each core has 64 KB L1 instruction cache, 96 KB L1 data cache, and 2MB L2 per core, while the chip has 164 MB of system level cache sharded across the chip.
One of the most revealing engineering choices is Olympus’ SMT-like behavior. SMT has been a staple in x86 since 2002 and can improve utilization by enabling two threads to use idle execution units in a cycle, with double-digit percentage gains for certain apps. Arm historically downplayed SMT in some of its agentic CPU approaches, but Nvidia implemented SMT-like functionality in Olympus, marketing it as “spatial multithreading.” The description differs from conventional SMT because it is not two threads sharing a single core’s resources the same way. Instead, Nvidia’s spatial multithreading gives two threads access patterns that look like core bifurcation, a design choice meant to extract more useful work without purely doubling throughput.
Zoom out and the platform stakes get real. Nvidia’s agentic AI reference designs envision cramming up to 128 of these superchips, meaning 256 CPUs, into a single liquid-cooled rack, totaling 22,528 cores and 384 TB of memory. That is not a subtle product launch. It is Nvidia testing whether hyperscalers will standardize on a CPU stack that is architecturally designed around AI agent execution and GPU orchestration, rather than simply buying another Arm license or letting Intel and AMD keep the server CPU mindshare.
For executives watching this space, the key second-order implication is that Nvidia is trying to own a new layer of the AI workflow. If Vera succeeds at becoming the head node and agent host, it changes how systems are built and who controls the timing, memory movement, and execution efficiency around GPUs. In a market where software bottlenecks are often the silent killer, Nvidia is betting that the CPU itself can stop being the weak link.
This story's Key Insights and Take-aways are locked.
Create a free account to unlock Executive Actions for one credit.
Register to UnlockAlways free for Executives Club members. Join the Club
More in Technology
Cyborg cockroaches can now carry cameras and inject medicine on command
A WIRED report shows electrodes, cameras, and injection devices turning live roaches into remote medics for disaster rescue.
Isar Aerospace's Spectrum reaches orbit on second flight, a European commercial first
The German startup's second-flight success lands days before Macron's Paris summit, giving Europe a homegrown launch option as SpaceX and Blue Origin bow out.
Tesla's wheel-less Cybercab rolls into China as sales stall
The EV maker will debut its autonomous robotaxi in Beijing and Shanghai mid-September, hoping its tech wow-factor reignites demand in its second-largest market.



