Nvidia shifts focus: CPUs dominate agentic data centers
Nvidia’s CPU Ambition: How a Monolithic Design is Redefining Data Center Power
In the high-stakes arena of artificial intelligence, the focus has been laser-sharp on graphics cards, but beneath the surface of the data center buildout, a revolution in central processing units is quietly underway. Nvidia, the undisputed giant of silicon Valley, is not resting on its laurels; it is making a monumental move into the CPU market with the Vera architecture, signaling a strategic shift that could redefine how massive AI operations are executed.
Ian Buck, Nvidia’s vice president of hyperscale and high-performance computing and the visionary behind CUDA, recently highlighted this trajectory. He noted that the company has shipped hundreds of thousands of Grace standalone servers and suggested that this deployment scale may be even larger, underscoring Nvidia‘s ambition to lead in a market increasingly defined by agentic AI workloads.
The demand for specialized processing has fundamentally changed hardware requirements. Where previous AI training often relied on configurations involving numerous GPUs per CPU, the new era of agentic AI is pushing systems toward more efficient one-to-one ratios. This shift created a hunger for CPUs built not just for raw power, but for unparalleled internal communication and efficiency—a perfect challenge for Nvidia’s next generation of processors.
Grace represented an initial on-ramp for Nvidia into the data center CPU space, utilizing Arm Neoverse V2 cores paired with Nvidia’s Scalable Coherency Fabric (SCF). Now, the Vera CPU represents a significant escalation. It incorporates an updated SCF alongside Nvidia’s first custom core design, Olympus, marking Nvidia‘s major entrance against established rivals like AMD and Intel.
What makes the Vera architecture truly distinctive is its monolithic design. Unlike many competitors who have favored chiplet designs to maximize core density—often trading off latency and coherency issues—Vera places all its 88 cores onto a single piece of silicon. This architectural choice allows Nvidia to dedicate substantial die space to the fabric itself, creating an incredibly robust system.
This commitment to internal coherence translates directly into performance metrics. Buck explained that this centralized approach provides a staggering 3.4 TB/s of core-to-core bandwidth within the CPU, ensuring every core can communicate with every cache and memory controller at full speed without bottlenecks. This massive fabric is the engine behind the system’s unified performance.
While chiplet designs offer flexibility in density, Nvidia‘s approach prioritizes cohesive performance. Buck acknowledged that this monolithic design involves a calculated trade-off: while it excels in high-speed internal communication, they recognize that some legacy data center workloads may still favor alternative approaches over the next few years.
Despite this acknowledgment, Nvidia’s vision remains expansive. The company forecasts that CPUs represent a massive Total Addressable Market (TAM) opportunity of $200 billion, a far more optimistic figure than industry averages by 2030. As agentic AI continues to expand the data center landscape—with estimates suggesting agents could add up to $60 billion to this market—Vera is positioned not just as a CPU, but as an integral component of the entire Nvidia ecosystem.
Vera is rapidly moving into full production, integrating seamlessly with other next-generation AI infrastructure, including Rubin GPUs and specialized networking components. This holistic approach ensures that Nvidia is building not just processors, but complete, high-performance computing environments for the future.