Nvidia maximizes compute from fixed power budgets Hot Chips 2026


Featured image Nvidia maximizes compute from fixed power budgets Hot Chips 2026

The race to build the next generation of AI data centers is no longer just about raw processing power; it’s a high-stakes battle over power. No matter how brilliant a single chip or GPU is, its ultimate performance is tethered by the physical reality of the facility—specifically, how much power can be delivered to the racks containing it. For future data center architects, the most critical constraint isn’t the silicon; it’s the power grid and the facility management.

This fundamental challenge is driving innovation, particularly in the realm of AI hardware. When Nvidia showcased the Rubin GPU architecture, they emphasized that unlocking the true potential of these systems demands sophisticated solutions for power management. They demonstrated that simply calculating peak power draw is an outdated approach; maximizing productivity requires a dynamic, intelligent approach to power allocation.

Enter Nvidia’s DSX MaxLPS suite—a system designed to overcome the limitations of static power provisioning. Traditional methods rely on conservative estimates of power consumption, often leading to inflated budgets and wasted energy because hardware rarely operates at peak capacity. This static planning creates rigidity: if one rack is idle while another is under heavy load, the system cannot dynamically reroute energy, leaving power stranded.

MaxLPS changes this equation by continuously monitoring power usage at the chip, rack, and group level. This intelligent control loop can identify unused power and redistribute it instantly to systems that need it most. The goal is simple: maximize the number of systems that can be installed and ensure superior performance per watt across the entire facility.

Nvidia’s vision allows data center operators to achieve remarkable efficiency. For example, in a measured scenario involving current GB300 racks, static provisioning might leave significant power unused due to utilization differences. However, dynamic management enables operators to safely install more hardware within the same power budget by exploiting these unused reserves.

This dynamic approach extends beyond just power distribution. The system also incorporates workload-specific power profiles, allowing the hardware to operate in modes optimized for inference, training, or general compute tasks. This mirrors the familiar experience of fine-tuning settings on a personal computer, but applied at a massive, infrastructure scale.

Furthermore, the Rubin systems incorporate advanced cooling techniques. By utilizing higher liquid coolant temperatures, such as 45 degrees Celsius, the racks become more thermally efficient. This shift helps reduce the energy demand of mechanical chillers—which can consume up to 40% of the site power budget—thereby significantly improving the facility’s power usage effectiveness (PUE).

Ultimately, by coupling dynamic power allocation with advanced thermal management, the industry gains flexibility over the lifespan of the installation. A facility can evolve its power strategy as AI workloads shift from intense training to efficient inference, allowing operators to free up capacity and unlock more profitable performance per watt.

Nvidia is providing the essential building blocks—the DSX toolkit—to help data center constructors and operators plan the future of AI infrastructure, ensuring that the power behind the revolution is as smart and efficient as the technology itself.

Image credit: Nvidia
Image credit: Nvidia
Image credit: Nvidia
Image credit: Nvidia

You may also like: