Subscribe
18:00Tao calls OpenAI’s Navier–Stokes push “resource extraction”17:10LAPTOP memecoin hits $190.81, then loses 99% inside an hour17:05Hubinger puts the odds of AI killing everyone above 10%; a colleague resigns16:39CancerBench launches; five frontier models tied at zero cancer types cured16:30ElevenLabs preparing 2028 IPO after $11bn round, The Information reports16:30Anthropic retracts its July explanation: Mythos 5 attacked systems knowingly
Hardware3 min read

Nvidia's first CPU is shipping, and its rack is sized at one agent sandbox per core

Vera has 88 Olympus cores and 1.2TB/s of memory bandwidth. The 256-CPU rack advertises 22.5K concurrent environments, which is 22,528 — the core count, exactly.

In briefVera has 88 custom Olympus cores, 1.2TB/s of memory bandwidth and up to 1.8x faster per-core performance on agentic AI workloads, and is shipping at scale1Ian Buck hand-delivered Vera systems to AWS, and previously to Oracle Cloud Infrastructure, Anthropic, OpenAI and SpaceXAI2Ian Buck's quote on the CPU moment in the AI factory3
An Nvidia board carrying Grace and Blackwell chips
Photo: 极客湾Geekerwan (CC BY 3.0)

Nvidia's Vera CPU is shipping at scale, with 88 custom Olympus cores, 1.2TB/s of LPDDR5X bandwidth and up to 1.5TB of capacity per socket. Ian Buck hand-delivered the first units to Anthropic, OpenAI, SpaceXAI and Oracle Cloud Infrastructure in May, and to AWS this week. The company published the delivery photographs on August 27th.

Skip the photographs. The rack arithmetic is the news.

Nvidia's Vera CPU Rack takes up to 256 Vera CPUs and is advertised as running "over 22.5K concurrent environments." Multiply 256 by 88 and you get 22,528.

One sandbox per core, rounded down for the marketing copy. Nvidia has decided the unit of agentic compute is a core rather than a container or a VM, and has priced a rack accordingly.

Now compare it with Grace, which is the fair comparison and the one Nvidia avoids making. Grace has 72 Arm Neoverse V2 cores and 500GB/s of LPDDR5X. Vera has 22% more cores and 2.4× the bandwidth.

Roughly double the memory bandwidth per core, in other words.

Memory bandwidth per core (GB/s)
Leading x86 with DDR5 (from NVIDIA's 3x claim)4.5NVIDIA Grace6.9NVIDIA Vera13.6

Why would a CPU need that? Because the work an agent generates is not floating point. Agents compile generated code, run Python toolchains, hold long-context state, move KV cache, and start and kill thousands of short-lived sandboxes — small, branchy, latency-bound work that a core-count-optimised server chip handles badly. Buck's framing is that this is "a new CPU moment in the AI factory," and on the evidence of what the rack is sold to do, he means it fairly literally.

Be careful with the headline speed-up, though. Nvidia's blog says Vera "delivers up to 1.8x faster per-core performance on agentic AI workloads." Nvidia's own product page says the same 1.8× applies to three named workloads "over leading x86 CPUs," with a footnote that the figure is "relative performance based on measured data, and subject to change." So it is a claim against Intel and AMD, not against Grace, and 1.8× is a ceiling rather than an average. Nvidia leaves the Grace comparison for you to compute. We did, above.

There is a strong argument against all of this: hyperscalers have been building their own Arm server CPUs for years and are not short of them. Oracle's answer is the one that matters here.

OCI plans to deploy hundreds of thousands of NVIDIA Vera CPUs beginning in 2026 because agentic AI demands sustained performance at massive scale. Vera's architecture is purpose-built for high-throughput reasoning workloads, delivering the efficiency, density and footprint OCI needs to power the next generation of enterprise AI.
Karan Batta, OCI, quoted by NVIDIA, Aug. 27, 2026

Nobody buys hundreds of thousands of anything for a host-processor role. So that is a standalone fleet, ordered as one.

Our read is that Vera's volume will come from the rack, not from the socket next to a Rubin GPU, and that reinforcement learning environments are the workload driving it — Anthropic's James Bradbury calls Vera "a promising part of the ecosystem when solving for agentic workloads," which is the sound of a compute lead who has already modelled it. We would also expect Nvidia never to publish a CPU revenue line, for the same reason it stopped publishing a networking one.

But hold onto one number until GTC. 22,528. If the next Vera rack advertises a sandbox count that is no longer a multiple of the core count, something in the software layer has changed.

Sources

01
Vera has 88 custom Olympus cores, 1.2TB/s of memory bandwidth and up to 1.8x faster per-core performance on agentic AI workloads, and is shipping at scaleVera packs 88 custom NVIDIA-designed Olympus cores, 1.2TB/s of memory bandwidth and delivers up to 1.8x faster per-core performance on agentic AI workloads. ... Core specs — 88 custom Olympus cores, 1.2TB/s memory bandwidth, up to 1.8x…” — blogs.nvidia.com · primary · Sep 10
02
Ian Buck hand-delivered Vera systems to AWS, and previously to Oracle Cloud Infrastructure, Anthropic, OpenAI and SpaceXAIAWS has received its first NVIDIA Vera CPU server and Vera Rubin GPU, hand-delivered in Seattle by NVIDIA Vice President of Hyperscale and HPC Ian Buck. AWS is the latest stop in Vera's rapid journey across the AI ecosystem. Previously,…” — blogs.nvidia.com · primary · Sep 10
03
Ian Buck's quote on the CPU moment in the AI factory"Agentic AI is creating a new CPU moment in the AI factory — as models move from answering to acting, Vera is purpose-built to keep that work moving at scale," Buck said.” — blogs.nvidia.com · primary · Sep 10
Show all 10 sources
04
The NVIDIA Vera CPU Rack integrates up to 256 Vera CPUs to run over 22.5K concurrent environmentsThe NVIDIA Vera CPU Rack powers reinforcement learning and agentic AI at AI factory scale. Built on NVIDIA MGX™, it integrates up to 256 Vera CPUs to run over 22.5K concurrent environments.” — nvidia.com · primary · Sep 10
05
Vera has 88 Olympus cores with Spatial Multithreading creating 176 threads, up to 1.2TB/s LPDDR5X and up to 1.5TB of memory, on a single compute die with 3.4TB/s SCF bisection bandwidthNVIDIA Vera features 88 custom Olympus cores built for the control-heavy, latency-sensitive work behind agentic AI and reinforcement learning. High single-thread performance helps software environments, tool calls, and evaluation loops…” — nvidia.com · primary · Sep 10
06
NVIDIA's product page states the 1.8x is over leading x86 CPUs across three workloads, and that Vera delivers 3x the bandwidth per core of leading x86 CPUs with DDR5NVIDIA Vera accelerates all three workloads by up to 1.8x over leading x86 CPUs, turbocharging the agentic inner loop to maximize AI factory output. Relative performance based on measured data, and subject to change. NVIDIA Vera CPU with…” — nvidia.com · primary · Sep 10
07
NVIDIA Grace has 72 Arm Neoverse V2 cores and up to 500GB/s of LPDDR5X bandwidth with 3.2TB/s SCF bisection bandwidthThe Grace CPU combines 72 high-performance and power-efficient Arm® Neoverse™ V2 cores, connected with the NVIDIA Scalable Coherency Fabric (SCF) that delivers 3.2TB/s of bisection bandwidth ... Grace is the first data center CPU to…” — nvidia.com · primary · Sep 10
08
Oracle plans to deploy hundreds of thousands of Vera CPUs beginning in 2026"OCI plans to deploy hundreds of thousands of NVIDIA Vera CPUs beginning in 2026 because agentic AI demands sustained performance at massive scale," Batta said. "Vera's architecture is purpose-built for high-throughput reasoning…” — blogs.nvidia.com · primary · Sep 10
09
Anthropic's head of compute James Bradbury's quote on Vera"Scaling compute is an important accelerant for the growth of models," Bradbury said. "We're excited to see Vera emerge as a promising part of the ecosystem when solving for agentic workloads."” — blogs.nvidia.com · primary · Sep 10
10
Vera handles orchestration, tool-calling, RL workloads, data analytics, agent sandboxing and long-context state managementWhat it handles — Orchestration, tool-calling, RL workloads, data analytics, agent sandboxing, long-context state management” — blogs.nvidia.com · primary · Sep 10
Up next · Keep readingHardware · 3 min read

A20 Pro is the first 2nm phone chip. The number that matters is 32 Neural Engine cores after three years at 16

Apple doubled the Neural Engine, widened memory bandwidth 50 percent, moved the DRAM off the thermal path and tripled the vapor chamber. Every one of those decisions is about running models on the phone.

Continue ↓