Subscribe
18:00Tao calls OpenAI’s Navier–Stokes push “resource extraction”17:10LAPTOP memecoin hits $190.81, then loses 99% inside an hour17:05Hubinger puts the odds of AI killing everyone above 10%; a colleague resigns16:39CancerBench launches; five frontier models tied at zero cancer types cured16:30ElevenLabs preparing 2028 IPO after $11bn round, The Information reports16:30Anthropic retracts its July explanation: Mythos 5 attacked systems knowingly
Hardware3 min read

Arm doubled its licensee compute subsystem and named two customers for its own CPU on the same day

Neoverse CSS N4 goes to 128 cores on TSMC N3P with PCIe Gen7 and FP8. Arm also said Oracle and ByteDance will deploy the AGI chip it builds itself.

In briefArm Neoverse CSS N4 supports 8 to 128 Neoverse N4 cores per die on TSMC N3P at up to 3.8GHz, with Armv9.3 as the baseline and FP8 and MMLA support1CSS N4 uses an Arm Neoverse CMN S4 interconnect supporting up to a 16x16 mesh and coherent multi-chiplet connectivity over CHI C2C2CSS N4 supports LPDDR6, DDR5 and MRDIMMs at 8000-12000 MT/s, and up to 128 lanes of PCIe Gen73
A processor die under magnification
Photo: cole8888 (CC BY-SA 2.0)

Arm launched Neoverse CSS N4 on September 8th: 8 to 128 Neoverse N4 cores on a single die built for TSMC N3P, up to 3.8GHz, 256MB of L3 per die, 2MB of L2 per core, DDR5 or LPDDR6, and up to 128 lanes of PCIe Gen7 with CXL 4.0. Armv9.3 becomes the baseline, and the cores add FP8 and MMLA.

Against the N2 subsystem it replaces, almost everything doubles or better. N2 stopped at 64 cores, 1MB of L2 per core, 64MB of L3, and 64 lanes of PCIe 5.0.

What Arm claims for CSS N4 versus Neoverse N3 (multiple)
Socket performance2Memory bandwidth1.75Performance per watt1.25

Those figures come with a configuration attached — 128 cores at 3GHz with 2MB of L2 each — and Arm has not said what clocks a 128-core part actually holds. Nor has it named a single partner. Which is normal for CSS launches, and not much of a risk: a handful of large contracts is the whole business.

Read the I/O list and you can probably date the product yourself. PCIe Gen7 and CXL 4.0 are not going into anything shipping next year. ServeTheHome makes the same point about the memory support, which stretches to MRDIMMs at 8000 to 12000 MT/s, and the interconnect is a CMN S4 mesh scaling to 16×16 with coherent multi-chiplet links over CHI C2C and UCIe. So this is likely a subsystem for silicon that tapes out in 2027 and appears in 2028 (which is also roughly when LPDDR6 volume arrives).

And now the awkward part.

On the same day, Arm said Oracle and ByteDance will deploy its own AGI CPU — a dual-die part with up to 136 Neoverse V3 cores, 272MB of L3 and up to 6TB of memory — alongside Meta, Lenovo, SAP, OpenAI and Cloudflare. Arm shipped an IP product for the companies that build server CPUs, and a customer list for the server CPU it builds itself, in one press cycle.

We would not call that a betrayal. Arm has been open about AGI serving as a demonstration vehicle, and its V-series is still under everyone's biggest chips: AWS Graviton, Google Axion, Azure Cobalt 200. The tension is quieter than a land grab, and it seems to show up in what CSS N4 is optimised for. N-series parts go into DPUs, IPUs and cost-sensitive cloud cores. Intel's IPU Adapter E2100 runs Neoverse N1. XSight Labs built a 64-core 800G DPU on CSS N2, which is exactly the case for CSS: if you are a small team, Arm has already done the base design. The V-series seat, the one Arm now occupies with its own product, is where the margin is.

The better argument against worrying is Nvidia, which points the other way entirely. Grace used Neoverse V2. Vera, announced two weeks ago, uses Nvidia's own Olympus cores. Arm's largest and most visible licensee has already stopped buying Arm's cores while keeping Arm's architecture (the licence that matters, and the one nobody gives up). A company losing core sales to its own customers has an obvious reason to sell a finished chip.

So which way does this go? Our read is that the architecture licence outlives the core licence, that CSS is Arm's answer for everyone who cannot afford a custom core team, and that this is a smaller and more durable business than designing CPUs against Oracle's other suppliers. We would expect the first named CSS N4 partner to be a networking or DPU vendor rather than a general-purpose server CPU, and to be announced before the middle of 2027.

One number Arm has not published for AGI, more than a year after first talking about it, is a benchmark. "More than 2x the performance per rack compared to the latest x86 systems" is an internal estimate. If you are buying, ask for the SPEC run.

Sources

01
Arm Neoverse CSS N4 supports 8 to 128 Neoverse N4 cores per die on TSMC N3P at up to 3.8GHz, with Armv9.3 as the baseline and FP8 and MMLA supportThe new Arm Neoverse N4 CSS is built for 3nm process and supports 8 to 128 cores per die. With Neoverse N4 cores, Armv9.3 architecture becomes the baseline. These cores are designed to operate at up to 3.8GHz given their focus on…” — servethehome.com · reported · Sep 10
02
CSS N4 uses an Arm Neoverse CMN S4 interconnect supporting up to a 16x16 mesh and coherent multi-chiplet connectivity over CHI C2CFor the core interconnect, there is an Arm Neoverse CMN S4 Interconnect as the fabric for the CSS N4 subsystem. CMN S4 supports up to a 16×16 mesh interconnect and a fully coherent multi-chiplet connectivity using CHI C2C.” — servethehome.com · reported · Sep 10
03
CSS N4 supports LPDDR6, DDR5 and MRDIMMs at 8000-12000 MT/s, and up to 128 lanes of PCIe Gen7On the memory side, this new generation also brings support for both LPDDR6 memory as well as DDR5 and MRDIMMs at 8000-12000MT/s speeds. For I/O, there are up to 128 lanes of PCIe Gen7 which also tells us that this is designed for CPUs…” — servethehome.com · reported · Sep 10
Show all 12 sources
04
CSS N4 offers up to 256MB of L3 per die, 2MB of L2 per core, and 128 lanes of PCIe 7/6 with CXL 4.0, versus CSS N2's 64 cores, 1MB L2 per core, 64MB L3 and 64 PCIe 5.0 lanesThe platform supports either DDR5 or LPDDR6, and features up to 256 MB of L3 cache per die. For local cache, Arm includes up to 2 MB of L2 per core, as well as 64 KB of L1 instruction cache and 64 KB of L1 data cache per core. For I/O,…” — tomshardware.com · reported · Sep 10
05
Arm claims CSS N4 at 128 cores and 3GHz delivers twice the socket performance of Neoverse N3, 1.25x performance per watt and 1.75x the memory bandwidthWith 128 cores running at 3GHz and 2MB of L2 cache per core, Arm says Neoverse CSS N4 delivers twice the socket performance of Neoverse N3, 1.25x performance per watt, and 1.75x the memory bandwidth.” — tomshardware.com · reported · Sep 10
06
Arm has not announced any CSS N4 partners, and clocks at maximum core count were not disclosedPresumably, the clocks drop as the core count rises; Arm didn't clarify the maximum clocks for each possible configuration. ... Arm has yet to announce any partners, though traditionally, only a few large CSS contracts are needed.” — tomshardware.com · reported · Sep 10
07
Arm revealed that Oracle and ByteDance will deploy its own AGI CPU alongside Meta, Lenovo, SAP, OpenAI and CloudflareAlongside the announcement of Arm Neoverse CSS N4, the company revealed additional deployments of its own AGI chip , which is built with Neoverse V3 cores. The company revealed that Oracle and ByteDance will deploy AGI chips, alongside…” — tomshardware.com · reported · Sep 10
08
Arm's AGI CPU is a dual-die part with up to 136 Neoverse V3 cores, 272MB of L3, up to 3.7GHz and up to 6TB of memory, and its performance claim is an internal estimateAGI is a dual-die CPU with up to 136 Neoverse V3 cores and up to 272 MB of L3 cache that can clock up to 3.7 GHz. It has the specs to match any high-end x86 design currently on the market, built on a 3nm node and packing up to 6TB of…” — tomshardware.com · reported · Sep 10
09
N-series cores are typically deployed in DPUs, IPUs and cost-sensitive cloud CPUs, including Intel's IPU Adapter E2100 on Neoverse N1 and Azure Cobalt 100 on Neoverse N2N-series cores aren't usually deployed in high-performance CPUs. Rather, they fit into less-performant accelerators, such as Intel's IPU Adapter E2100, which is built on Neoverse N1 cores. We've also seen it deployed in less-demanding,…” — tomshardware.com · reported · Sep 10
10
XSight Labs' E1 64-core 800G DPU is built on an Arm Neoverse CSS N2 designOne of the neatest N-series CSS implementations we have seen recently is the XSight Labs E1 64-Core Arm 800G DPU . This is an Arm Neoverse CSS N2 design with 64 Arm Neoverse N2 cores. ... For a startup like XSight Labs, using a Neoverse…” — servethehome.com · reported · Sep 10
11
Nvidia's Grace CPU used Arm Neoverse V2 cores while Vera uses Nvidia's own Olympus coresThe Grace CPU combines 72 high-performance and power-efficient Arm® Neoverse™ V2 cores, connected with the NVIDIA Scalable Coherency Fabric (SCF)” — nvidia.com · primary · Sep 10
12
Vera uses 88 custom NVIDIA-designed Olympus coresVera packs 88 custom NVIDIA-designed Olympus cores, 1.2TB/s of memory bandwidth and delivers up to 1.8x faster per-core performance on agentic AI workloads.” — blogs.nvidia.com · primary · Sep 10
Up next · Keep readingHardware · 3 min read

A20 Pro is the first 2nm phone chip. The number that matters is 32 Neural Engine cores after three years at 16

Apple doubled the Neural Engine, widened memory bandwidth 50 percent, moved the DRAM off the thermal path and tripled the vapor chamber. Every one of those decisions is about running models on the phone.

Continue ↓