Tag: Custom Chip

  • Meta’s solution to the global memory shortage is to use DDR4 in a DDR5 server, with a custom chip making the impossible possible

    Featured image Metas solution to the global memory shortage is to use DDR4 in a DDR5 server with a custom chip making the impossible possible

    Imagine being a titan of technology, like Meta, spending millions on cutting-edge servers that boast blazing fast DDR5 memory. Yet, despite this incredible speed, you still hit a wall: not enough memory. The solution isn’t just bigger chips; it’s a clever trick involving taking the old, slower DDR4 modules and weaving them back into the system to create an enormous pool of RAM.

    This isn’t science fiction; it’s emerging reality. As CXL technology, coupled with custom hardware, is bridging the gap between modern CPUs and legacy memory standards. The breakthrough lies in finding a way for the DDR4 modules to speak the language of the DDR5 processors—a task managed by specialized components like Meta‘s custom chip, Vistara.

    How does this magic happen? CXL acts as the translator, managing the complex relationship between the two memory types. The Vistara chip then orchestrates the entire operation, treating the slower DDR4 modules not as primary storage, but as a vast reservoir of ‘cold storage.’ This allows the system to intelligently manage data: actively used information is kept in the blazing-fast DDR5 ‘hot storage,’ while less immediate or background data resides in the accessible, yet slower, DDR4 pool.

    This innovative approach addresses one of the biggest bottlenecks in modern computing. By strategically partitioning memory, systems can maximize performance without requiring prohibitively expensive, monolithic amounts of high-speed RAM. It’s about optimizing latency and bandwidth to achieve peak efficiency.

    Consider Meta‘s implementation of this concept. A single MemServer node, featuring powerful processors like the AMD Epyc 9000-series chips running DDR5 memory, is augmented by these Vistara expansion cards. This setup effectively combines high-speed performance with massive capacity, allowing one server unit to achieve a grand total of 1,024 GB, or 1 TB, of system memory—more than enough capacity to handle massive simulations like Star Citizen.

    The potential extends beyond massive data centers and into consumer gaming. While integrating this architecture into standard desktop PCs faces hurdles related to the universal adoption of CXL standards and CPU design, the underlying principle remains compelling. If DRAM prices continue their trajectory toward normalcy, this technological evolution could offer a practical workaround for the enormous memory demands in high-end gaming rigs.

    Ultimately, the future of computing isn’t just about raw speed; it’s about intelligent resource management. By leveraging older technologies through innovative custom solutions, engineers are finding clever ways to unlock unprecedented levels of performance and capacity.

  • Broadcom and OpenAI unveil custom-built Jalapeño inference processor — OpenAI’s first chip is a massive reticle-sized ASIC built in an ultra-fast nine-month development cycle

    The race for AI supremacy isn’t just about bigger models; it’s about smarter hardware. This is where OpenAI and Broadcom are throwing down a serious challenge with Jalapeño, a custom-built inference processor designed from the ground up to power the next generation of large language models and agentic AI workloads.

    Jalapeño is not just another AI accelerator. It is presented as a purpose-built inference ASIC, meticulously engineered around the specific behaviors of LLMs. Rather than repurposing existing training hardware, the architecture addresses fundamental bottlenecks that plague large-scale AI: inefficient data movement, the crucial balance between compute and memory resources, networking efficiency, and overall operational behavior.

    At the heart of Jalapeño’s design is a focus on maximizing both throughput and minimizing latency. To achieve this, the team opted for a sophisticated configuration, utilizing a massive compute chiplet surrounded by six high-bandwidth HBM memory modules. This design choice allows the processor to execute demanding reasoning and agentic tasks with exceptional efficiency.

    The promise of Jalapeño lies in its efficiency. The companies claim that this optimized architecture delivers performance per watt substantially higher than current state-of-the-art hardware, suggesting a dramatic leap in energy efficiency for complex AI calculations. While specific benchmarks remain under wraps, internal testing indicates that the chip is efficiently executing cutting-edge workloads, including GPT-5.3-Codex-Spark.

    The physical reality of Jalapeño underscores its ambition. Engineering samples are already operating successfully in the lab, and the development cycle itself was remarkably swift, reaching tape-out in just nine months. This acceleration was made possible by integrating artificial intelligence into the chip design process, alongside Broadcom’s established practice of reusing logic across multiple custom designs.

    The design philosophy extends beyond immediate use. Jalapeño is envisioned to support not only OpenAI’s initiatives but also the broader industry of LLMs, potentially positioning it as a platform that others can leverage. The goal is to create a scalable physical infrastructure capable of handling the demands of the next decade of AI.

    Broadcom‘s CEO emphasized this vision, stating that the collaboration with OpenAI represents a fundamental commitment to scaling the physical infrastructure needed for future AI deployment. This hardware is slated for deployment in gigawatt-scale data centers starting in 2026, working alongside partners like Microsoft.

    The implications are vast. By prioritizing these holistic architectural considerations—from kernel execution to memory architecture—Jalapeño aims to set a new benchmark for how AI compute is delivered, pushing the boundaries of what modern silicon can achieve.