Fujitsu Monaka CPU stacks cache on separate 5nm die


Featured image Fujitsu Monaka CPU stacks cache on separate 5nm die

Fujitsu’s Monaka CPU: Redefining AI Performance Through 3D Architecture

The race to build truly efficient and powerful AI infrastructure is heating up, demanding chips that can deliver massive computational muscle while respecting strict power envelopes. Into this high-stakes environment, Fujitsu has unveiled the Monaka server CPU, a design that isn’t just incremental—it’s a fundamental rethinking of how silicon is stacked and optimized for green AI data centers.

At the heart of Monaka is a groundbreaking architectural strategy. Rather than traditional layouts, Fujitsu has employed a three-tiered approach, essentially stacking the components to maximize efficiency. The design features a 2nm core die built on TSMC N2P, which sits face-to-face with a separate 5nm SRAM die that houses the entire last-level cache. This separation allows the core to run hottest, while the crucial cache remains optimally placed, separated by a silicon interposer from the I/O die. This innovative setup dramatically changes the dynamics of data flow, separating the core compute from the massive cache, a distinction that positions Monaka uniquely against competitors like AMD’s 3D V-Cache.

This architectural ingenuity is coupled with a focus on future-proofing performance. While previous generations utilized 512-bit vector units, Monaka introduces dual 256-bit SVE2 vector units per core. This adjustment is not merely a reduction; it is a strategic choice, designed to minimize core size and SIMD width for general-purpose code, allowing Fujitsu to prioritize density and cost-effective performance for the data center.

Fujitsu’s approach extends beyond transistor arrangement and delves deep into energy efficiency. The chip was engineered specifically for power conservation, utilizing low-dropout voltage regulators placed strategically across the die to feed dynamic voltage and frequency scaling. This ultra-low-voltage operation allows the 144 cores to operate at roughly 30% below nominal voltage for half the power consumption, delivering energy savings comparable to moving one full generation beyond the 2-nanometer node.

Monaka is not just a piece of hardware; it is part of a larger national strategic effort. Developed with support from Japan’s New Energy and Industrial Technology Development Organization, the chip is explicitly aimed at fueling Japan’s green AI data centers. This commitment is reflected in the design choices, which prioritize power efficiency and performance metrics like DGEMM and INT8 throughput, promising up to two-times AI performance improvements compared to previous benchmarks.

The implications for the wider semiconductor landscape are significant. By focusing on cache-on-die stacking and high-bandwidth DDR5 connectivity, Monaka competes not just on core count, but on holistic system performance. Furthermore, the success of Monaka sets a compelling precedent for Japan’s semiconductor goals, especially given the ongoing commitment to domestic fabrication capacity through initiatives like Rapidus and the focus on building a dedicated 1.4nm AI chip entirely in Japan.

Looking ahead, Monaka paves the way for future supercomputing ambitions. Its successor, FugakuNEXT, is already being positioned as the core for Japan’s next flagship supercomputer, demonstrating how these cutting-edge designs can power the next generation of global research and AI development.

You may also like: