Huawei AI roadmap doubles Ascend NPU expectations
Huawei Charts a Bold New Course for AI Hardware with Accelerated Roadmap
The world of artificial intelligence hardware is undergoing a seismic shift, and at the forefront of this revolution is Huawei. At its recent annual Connect event, the company unveiled a radically updated AI hardware roadmap, signaling a major evolution in their approach to developing next-generation accelerators and supporting processors.
This update isn’t just about new chips; it reflects a deep architectural overhaul. Huawei is transitioning away from its previous SIMD architectures, which have served them well for nearly a decade, toward a sophisticated SIMD+SIMT framework. This new architecture merges vector-based processing with thread-level parallelism, aiming to dramatically improve hardware utilization and performance across diverse and complex AI workloads.
Underpinning this evolution is the introduction of the new Peerium architecture, built upon the UnifiedBus, which provides a foundation for greater flexibility and power. This strategic move underscores Huawei’s commitment to creating highly optimized silicon designed specifically for the demands of modern machine learning.
The transition is already yielding results. The initial products in this new paradigm, such as the Ascend 950PR for prefill and recommendation tasks and the Ascend 950DT for decoding and training, are already gaining traction. Systems like the Atlas 950 SuperPoD are entering large-scale commercial use, demonstrating the immediate viability of this new infrastructure.
Looking ahead, the timeline for the next generation of accelerators is tightly packed, with Huawei maintaining a focused, one-generation-per-year cadence. This accelerated schedule positions the company to deliver powerful hardware in a record time.
The Ascend 960-series is set to lead this acceleration, with the Ascend 960DT slated for formal availability in the first quarter of 2027. This accelerator is poised to deliver 2 FP8 PFLOPS and 4 FP4 PFLOPS, backed by 288 GB of memory and impressive interconnect speeds.
Following this, the Ascend 960PR NPU will arrive in the third quarter of 2027. What makes this model particularly exciting is its dramatically enhanced low-precision capability: it is expected to deliver 8 FP4 PFLOPS for inference—twice the performance compared to previous announcements—while still retaining strong performance for training.
The roadmap continues to climb. The Ascend 970, due in 2028, is expected to deliver 3.6 FP8 PFLOPS and 14 FP4 PFLOPS. Finally, the Ascend 980, anticipated in 2029, promises to push further boundaries, targeting 7.2 FP8 PFLOPS and 28 FP4 PFLOPS, along with massive bandwidth increases up to 38.4 TB/s. These figures, though preliminary, paint a picture of AI hardware that is rapidly evolving into the next generation.