Arm Ranger platform hits 128 cores per die


Featured image Arm Ranger platform hits 128 cores per die

Arm is not just designing the future of mobile processors; it is reshaping the entire landscape of high-performance computing by delivering flexible, powerful silicon platforms to the cloud. With the introduction of the Neoverse CSS N4 platforms, Arm is offering developers and hyperscalers a way to build highly customized chips that maximize performance per watt, moving beyond off-the-shelf limitations.

At the heart of this innovation is the Compute Subsystem, or CSS, which functions as a semi-custom program. This architecture allows customers to leverage Arm’s proven Intellectual Property (IP) blocks and configure every component—from core count and cache size to I/O and connectivity—to perfectly suit their specific demands. This is the same flexible approach seen across major cloud providers and chip designers, from Azure and Google Cloud to Nvidia and Intel.

The new Neoverse CSS N4 platform, built on TSMC’s N3P process, brings serious muscle. These systems can feature between eight and 128 Neoverse N4 cores, clocking up to 3.8 GHz. Crucially, the architecture is designed for massive scalability, supporting multi-chiplet and multi-socket designs, alongside advanced chip-to-chip interconnects like UCIe, giving architects unprecedented freedom in designing complex systems.

Powering this versatility are powerful metrics. The N4 architecture not only boosts raw capacity but dramatically improves efficiency. Arm reports that these systems deliver twice the socket performance compared to the Neoverse N3, along with 1.25x performance per watt and 1.75x memory bandwidth. This efficiency is key to squeezing maximum computational power out of every watt of energy.

The platform is loaded with memory and cache capabilities, featuring up to 256 MB of L3 cache per die, along with 2 MB of L2 cache per core. For I/O, the CSS N4 supports up to 128 lanes of PCIe 7/6 and CXL 4.0, ensuring that data can flow instantly and efficiently across the entire system.

Arm has strategically developed different core types to meet varied needs. The N-series cores are optimized for excellent performance per watt, making them ideal for general acceleration and cloud workloads, often fitting into accelerators like Intel’s IPU Adapter E2100. Meanwhile, the V-series cores are geared toward maximum raw performance, perfect for demanding CPU tasks.

This flexible approach has already seen massive deployment, particularly in the realm of Artificial General Intelligence (AGI). Arm’s own AGI chip, built on Neoverse V3 cores, is already being adopted by major players like Oracle and ByteDance, alongside deployments at Meta, Cloudflare, and OpenAI. The AGI architecture itself is a dual-die CPU with up to 136 cores and substantial memory capacity, designed to handle the demands of the most complex AI models.

Arm’s vision is clear: to be the foundational engine for custom silicon development across the industry. By providing validated building blocks, Arm allows partners to rapidly create bespoke processors tailored to the unique needs of their applications. As Arm continues to develop its roadmap, including the upcoming Vega architecture, the focus remains firmly on empowering the next generation of computing with unparalleled flexibility and efficiency.

You may also like: