AMD and Cerebras accelerate AI inference with WSE and EPYC
The quest for smarter, faster Artificial Intelligence demands more than just bigger chips—it requires fundamentally rethinking how we build computing systems. At the cutting edge of this revolution, AMD and Cerebras Systems are forging an ambitious new path by combining their specialized hardware into a unified platform designed to unleash unprecedented AI performance.
This collaboration centers on creating a disaggregated inference architecture, promising to shatter previous limitations by assigning specific tasks to architectures optimized for those jobs. The goal is simple: combine the low-latency prowess of AMD’s CPUs and Instinct GPUs with the massive throughput capabilities of Cerebras’ Wafer-Scale Engines (WSE).
The resulting system marries two distinct compute environments into one cohesive workflow. AMD’s Helios rack infrastructure, powered by EPYC CPUs and Instinct MI400-series accelerators, is tasked with handling the heavy lifting of prompt processing and managing vast context windows—the initial cognitive stage of an AI request.
Meanwhile, Cerebras’ WSE processors step in to tackle the memory-bandwidth-intensive token-generation stage. This division of labor ensures that each component performs its specialized function with maximum efficiency, leading to a system designed for peak operational performance.
This specialized approach is not new, but AMD and Cerebras are taking the established philosophy—similar to Nvidia’s concept—and inverting the roles. Where traditional designs separate inference into context/prefill and generation/decode stages, this new platform specializes them: the Helios system focuses on processing prompts, while the WSE focuses squarely on latency-sensitive token generation.
The payoff of this architectural specialization is substantial. By allowing different parts of the workload to run on dedicated hardware optimized for their specific demands, the combined AMD and Cerebras platform aims to deliver up to five times higher tokens per second per watt (T/s/W). This dramatic efficiency gain translates directly into more powerful and sustainable AI deployment.
The vision extends beyond just raw speed; it’s about creating a holistic ecosystem where compute capacity meets generation speed seamlessly. Cerebras plans to integrate AMD Helios systems into its own data centers, linking them directly with its WSE racks to realize this combined potential.
This integrated offering is scheduled to become available initially through Cerebras Cloud in the second half of 2026, signaling a major milestone in how high-performance AI infrastructure will be designed and deployed. It’s a blueprint for the future: specialized hardware working in perfect harmony to redefine the limits of what AI can achieve.