Nvidia's first CPU is shipping, and its rack is sized at one agent sandbox per core
Vera has 88 Olympus cores and 1.2TB/s of memory bandwidth. The 256-CPU rack advertises 22.5K concurrent environments, which is 22,528 — the core count, exactly.

Nvidia's Vera CPU is shipping at scale, with 88 custom Olympus cores, 1.2TB/s of LPDDR5X bandwidth and up to 1.5TB of capacity per socket. Ian Buck hand-delivered the first units to Anthropic, OpenAI, SpaceXAI and Oracle Cloud Infrastructure in May, and to AWS this week. The company published the delivery photographs on August 27th.
Skip the photographs. The rack arithmetic is the news.
Nvidia's Vera CPU Rack takes up to 256 Vera CPUs and is advertised as running "over 22.5K concurrent environments." Multiply 256 by 88 and you get 22,528.
One sandbox per core, rounded down for the marketing copy. Nvidia has decided the unit of agentic compute is a core rather than a container or a VM, and has priced a rack accordingly.
Now compare it with Grace, which is the fair comparison and the one Nvidia avoids making. Grace has 72 Arm Neoverse V2 cores and 500GB/s of LPDDR5X. Vera has 22% more cores and 2.4× the bandwidth.
Roughly double the memory bandwidth per core, in other words.
Why would a CPU need that? Because the work an agent generates is not floating point. Agents compile generated code, run Python toolchains, hold long-context state, move KV cache, and start and kill thousands of short-lived sandboxes — small, branchy, latency-bound work that a core-count-optimised server chip handles badly. Buck's framing is that this is "a new CPU moment in the AI factory," and on the evidence of what the rack is sold to do, he means it fairly literally.
Be careful with the headline speed-up, though. Nvidia's blog says Vera "delivers up to 1.8x faster per-core performance on agentic AI workloads." Nvidia's own product page says the same 1.8× applies to three named workloads "over leading x86 CPUs," with a footnote that the figure is "relative performance based on measured data, and subject to change." So it is a claim against Intel and AMD, not against Grace, and 1.8× is a ceiling rather than an average. Nvidia leaves the Grace comparison for you to compute. We did, above.
There is a strong argument against all of this: hyperscalers have been building their own Arm server CPUs for years and are not short of them. Oracle's answer is the one that matters here.
OCI plans to deploy hundreds of thousands of NVIDIA Vera CPUs beginning in 2026 because agentic AI demands sustained performance at massive scale. Vera's architecture is purpose-built for high-throughput reasoning workloads, delivering the efficiency, density and footprint OCI needs to power the next generation of enterprise AI.
Nobody buys hundreds of thousands of anything for a host-processor role. So that is a standalone fleet, ordered as one.
Our read is that Vera's volume will come from the rack, not from the socket next to a Rubin GPU, and that reinforcement learning environments are the workload driving it — Anthropic's James Bradbury calls Vera "a promising part of the ecosystem when solving for agentic workloads," which is the sound of a compute lead who has already modelled it. We would also expect Nvidia never to publish a CPU revenue line, for the same reason it stopped publishing a networking one.
But hold onto one number until GTC. 22,528. If the next Vera rack advertises a sandbox count that is no longer a multiple of the core count, something in the software layer has changed.
