Tag: US

  • China’s LineShine supercomputer dethrones US’ El Capitan, secures first place in Top 500 list — first machine in the rankings to sustain more than 2 ExaFLOPS of double-precision performance using only CPUs

    China’s LineShine Supercomputer Redefines the Limits of Computing

    A new titan has entered the world of High-Performance Computing (HPC), and it’s hailing from China. The LineShine supercomputer has successfully dethroned El Capitan to claim the top spot globally, proving that domestic technological prowess can compete with the world’s most advanced supercomputing systems.

    The achievement is backed by staggering raw power. In the Linpack benchmark, LineShine delivered an impressive 2.198 FP64 ExaFLOPS. What makes this feat even more remarkable is that it achieved this level of double-precision performance using only CPUs—establishing it as the first machine in the Top 500 list to sustain over 2 ExaFLOPS using solely CPU power.

    This formidable system is anchored at the National Supercomputing Centre in Shenzhen and represents a powerful collaboration between the Shenzhen Cloud Computing Center and the National Supercomputer Center in Shenzhen. It operates by harnessing 13.79 million cores, interconnected by a proprietary LingQi network, and demands 42.2 MW of energy to perform its calculations.

    Under the hood, LineShine utilizes semi-custom 304-core LX2 processors built on the Armv9 instruction set architecture, clocked at 1.55 GHz. Each core is packed with advanced features like Arm SVE and SME units, designed to accelerate vector and matrix operations essential for both scientific computing and emerging AI workloads, supporting various data formats including FP64, FP32, BF16, FP16, and INT8.

    The memory architecture is equally ambitious, pairing 32 GB of on-package HBM with up to 256 GB of external DDR5 memory, providing massive bandwidth of up to 4 TB/s to feed the hungry processors.

    While the sheer FP64 performance is astounding, an analysis of efficiency paints a more nuanced picture. LineShine delivered 52.07 GFLOPS/W. Although this figure sits just below El Capitan’s benchmark of 60.94 GFLOPS/W, it still demonstrates exceptional power management capabilities.

    Crucially, when compared to other CPU-only supercomputers like Fugaku—which struggles to deliver efficiency between 14.78 and 16.84 GFLOPS/W—LineShine clearly pulls ahead in terms of performance per watt. This efficiency gap highlights the sophisticated engineering behind the LineShine architecture.

    Despite its dominance in traditional supercomputer tasks, the system shows limitations when tackling modern mixed-precision AI workloads. In the HPL-MxP test, LineShine achieved 7.92 mixed-precision EFLOPS, placing it behind giants like El Capitan, Frontier, and Aurora. This suggests that while raw computational muscle is massive, incorporating specialized accelerators remains key for maximizing performance in contemporary AI training and inference.

    Nonetheless, the very fact that a supercomputer developed domestically has achieved such extraordinary FP64 performance is a landmark achievement. Furthermore, LineShine‘s submission to the Top 500 ranking signals confidence in its homegrown technologies, affirming a pathway for independent supercomputing development free from external technological constraints.