NVIDIA is shifting AI data center optimization away from isolated components and toward the joint management of power, cooling, networking, and workloads. Its NVIDIA DSX platform aims to increase the amount of AI work an infrastructure can handle within a fixed power budget, with Lambda tests showing 24% more tokens per second under the same power cap.
NVIDIA DSX in 30 seconds
- Lambda ran 19 nodes within the power budget previously reserved for 16 nodes at full power.
- Cluster throughput went from about 4 to 5 million tokens per second, a 24% increase.
- Performance per watt improved 23% in that validation.
- Emerald AI automatically scaled a load down from 4 to 3 megawatts in response to signals from Silicon Valley Power.
- NVIDIA projects up to 40% more GPU capacity in certain Vera Rubin deployments within the same power budget.
The growth of artificial intelligence is turning power availability into one of the constraints shaping the growth of accelerated computing infrastructure. NVIDIA argues that adding more megawatts isn’t always the only way to gain capacity: the electricity already available can also be put to better use.
That’s the thinking behind NVIDIA DSX, a platform that combines design and simulation tools with operating software, dynamic power management, and reference architectures. NVIDIA introduced DSX at GTC Taipei and is using its first validations to show how it can increase the work done per unit of energy.
The metric NVIDIA puts at the center of this strategy is the amount of useful work an AI factory can produce per megawatt. In inference-dedicated systems, this can be observed through tokens generated per second, although the final result depends on the workloads, the configuration, and the specific conditions at each site.
Lambda gets 24% more tokens with the same power budget
One of the most concrete tests comes from Lambda, a GPU infrastructure provider. The company validated NVIDIA DSX MaxLPS on a five-rack, 19-node cluster based on NVIDIA HGX B200 servers.
The system uses dynamic power software to monitor GPU and rack consumption and redistribute available headroom according to workload needs. The idea is to prevent a static configuration from reserving power that isn’t actually being used at a given moment.
In the test, Lambda managed to run 19 nodes within the same power budget it used as a reference for 16 nodes running at full power. The result was roughly a 24% increase in aggregate token throughput, from about 4.04 million to roughly 5 million tokens per second.
Performance per watt also rose 23%. NVIDIA and Lambda note that this behavior is especially relevant when training and inference share infrastructure, because their consumption profiles aren’t identical and can leave headroom that a static allocation system fails to capture.
NVIDIA’s documentation describes MaxLPS as a dynamic allocation model that aims to recover that available capacity within a set power limit. In its reference scenarios for future Vera Rubin NVL72-based factories, NVIDIA says the approach could allow up to 40% more GPUs within the same power budget in certain environments. That’s a company estimate, not a general result applicable to any data center.
The technical documentation itself gives an illustrative inference scenario with a 1-megawatt budget in which MaxLPS reaches a capacity of 400 GPUs and token throughput equivalent to 1.35 times MaxP mode. NVIDIA notes that realizing that potential requires designing the facility with cooling and other infrastructure elements in mind as well.
An AI factory can also respond to the power grid
The other part of the proposal concerns the relationship between data centers and the electrical grid. NVIDIA describes a demonstration carried out with Emerald AI and Silicon Valley Power, the utility that supplies the Santa Clara area.
In one of the episodes described by NVIDIA, Silicon Valley Power sent a signal to reduce an AI factory’s consumption. Emerald AI Conductor received the instruction and automatically adjusted flexible loads: lower-priority jobs were slowed down or rescheduled while tasks considered priority kept running.
Power dropped from 4 to 3 megawatts with no manual intervention. NVIDIA also notes that Silicon Valley Power has sent more than 200 demand signals to that facility and that the mechanism responded in every case the company described.
This demonstration uses Emerald AI Conductor and is not yet a DSX Flex installation. NVIDIA presents it as a prior experience that lays the groundwork for the DSX Flex strategy, which is designed to receive grid signals — such as load-reduction requests, demand-response events, or certain price signals — and adjust the priority of AI workloads.
According to NVIDIA, DSX Flex’s first specific commercial deployment is planned for a 96-megawatt facility in Manassas, Virginia, inside the company’s AI Factory Research Center. The platform will incorporate Emerald AI Conductor integration as DSX Flex matures.
The concept slightly changes how we understand an AI factory. The infrastructure doesn’t just consume electricity — it can temporarily adjust its demand when certain loads allow it. Priority tasks continue, while others can wait for a moment with less pressure on the grid.
800 VDC and a unified view of the infrastructure
NVIDIA is also working on an 800 VDC power distribution architecture — that is, 800-volt direct current. The company is pursuing this design to reduce conversions, improve power delivery efficiency, and make it easier to power increasingly dense accelerated-computing racks.
The 800 VDC architecture is part of DSX’s reference designs. NVIDIA also points to a projected 3% to 5% improvement in end-to-end efficiency compared with a 54V distribution in certain Vera Rubin NVL72-related scenarios, with availability expected in 2027. That figure is an NVIDIA projection and depends on the configuration used.
DSX’s full proposal includes several pieces. DSX MaxLPS focuses on getting more performance within a fixed power budget; DSX Flex aims to adapt workloads to grid conditions; DSX OS provides software for managing the factories’ lifecycle; DSX Sim allows facilities to be modeled before they’re built; and the reference designs bring together validated compute, networking, storage, and installation-system configurations.
NVIDIA also folds cooling into that approach. The company notes that certain GB200 NVL72 racks with direct liquid cooling need to dissipate around 120 kW of heat. That’s why increasing compute power without also considering cooling, networking, or power distribution at the same time can leave capacity underused.
NVIDIA’s thesis is that an AI factory’s efficiency should be measured by the work it manages to complete with the resources available, not just a GPU’s peak specs. Lambda’s results provide an initial practical validation of that idea, while the Emerald AI experience points to another possibility: that certain facilities can adjust their consumption to meet the grid’s needs.
The next step will be seeing how these techniques perform in larger deployments and with different workloads. Lambda’s figures correspond to one specific configuration, while the projected improvements for Vera Rubin depend on each facility’s conditions.
Frequently asked questions
What is NVIDIA DSX?
NVIDIA DSX is a platform for designing and operating AI factories that combines compute, power, cooling, networking, and software with the goal of increasing the performance obtained within a set energy budget.
How much did Lambda’s performance improve?
Lambda achieved 24% more aggregate token throughput in a test with DSX MaxLPS, going from about 4.04 million to roughly 5 million tokens per second within the same power budget.
What does DSX Flex do?
DSX Flex is designed to adapt an AI factory’s workloads to power grid signals, temporarily reducing consumption for certain tasks while keeping priority loads running.
What role does the 800 VDC architecture play?
NVIDIA is incorporating 800 VDC into its reference designs to reduce conversion complexity and improve power distribution in facilities with increasingly dense compute racks.
via: blogs.nvidia

