NVIDIA is shifting the artificial intelligence race away from how many GPUs are installed and toward a metric that matters more and more: how much AI work a data center can produce with each available megawatt. Its answer is DSX MaxLPS, a set of power-management, software, liquid-cooling, and facility-design technologies with which the company projects that future Vera Rubin NVL72-based data centers could house up to 40% more GPUs within the same power budget.
NVIDIA DSX MaxLPS in 30 seconds
- MaxLPS aims to make better use of power that, at many data centers, is already the resource limiting new GPU deployments.
- NVIDIA proposes dynamically reallocating the electrical capacity that racks aren’t currently using.
- In representative tests, it holds performance steady while reducing the power provisioned per rack.
- Vera Rubin NVL72 is designed for liquid cooling with inlet water at up to 45°C.
- NVIDIA estimates up to 40% more Rubin GPUs within a facility with the same available power.
The approach matters because it isn’t promising that 40% by increasing the power contracted from the grid. The proposal is precisely to work within a fixed limit and recover capacity that’s currently lost to cooling, electrical distribution, operational reserves, and the gap between the maximum power provisioned and the racks’ actual consumption.
NVIDIA calls this MaxLPS, short for Maximum Land Power Shell — a way of treating three physical constraints of a facility together: available land, electrical power, and the building itself with its power, cooling, communications, and compute systems. For the company, these resources should be managed as a complete system rather than sizing each rack as an isolated unit.
The Problem of Reserving Electricity That Ends Up Going Unused
A data center can’t assume its servers will always draw their average power.
The electrical infrastructure has to handle peaks. If a rack can reach 120 kW under certain loads, reserving only 90 kW because that’s its typical draw could cause problems once the system hits its maximum.
The result is that each rack is traditionally allocated enough power to cover those scenarios.
But AI workloads don’t draw power at a constant rate.
Training, inference, synchronization, memory access, communications, checkpointing, prefill, and decode all generate different power profiles. There are moments when the GPUs are working intensely and others when they’re waiting on data or communications.
A rack may have power reserved that it isn’t temporarily using, while another rack could make use of it.
NVIDIA wants to turn that margin into shared capacity.
Its example starts from a hypothetical 100 MW facility. Under the model the company uses, 20 MW go to running the facility itself, 10 MW are lost in rack-level infrastructure, and another 10 MW fall outside useful AI workload due to operational inefficiencies tied to failures, restarts, and checkpoints.
That leaves roughly 60 MW going directly to the AI workload.
That breakdown shouldn’t be read as a universal rule for every data center. NVIDIA presents it as an illustrative scenario to explain how much the electrical power a facility receives can diverge from the power that ultimately produces computational work.
That raises an issue that matters more and more for operators: buying a more efficient GPU can help, but so can making sure a larger share of the contracted megawatts actually reaches those GPUs.
From a Fixed Power Budget per Rack to Dynamic Allocation
One of the pieces of MaxLPS is Dynamic Power Software (DPS), currently presented by NVIDIA as a Developer Preview.
The system monitors consumption from the facility level down to resource groups, racks, nodes, and GPUs. Operators set limits and policies, and the software continuously compares the power allocated against what the hardware is actually drawing.
When it finds available capacity, it can reallocate it within the managed group.
The idea resembles other dynamic-allocation mechanisms used for years in computing: reserve enough resources to guarantee operation, but take statistical advantage of the fact that not every component hits its maximum at the same time. The difference is that here the shared resource is kilowatts.
NVIDIA shows a scenario in which static provisioning leaves 170 kW unused within a total budget of 540 kW. Dynamic allocation would recover that margin and allow another rack to be added without expanding the facility’s total power.
The overall electrical limit doesn’t change.
No new electricity appears, either.
What changes is how the existing capacity gets used.
This nuance matters because the system can’t eliminate the data center’s physical limits. If every rack simultaneously needs its maximum power, there’s no margin left to redistribute. The benefit will depend on how the workloads actually behave and on how well their peaks can complement one another.
GB200 and Vera Rubin Show How Much Density Can Shift
NVIDIA provides results from representative inference workloads to illustrate the effect.
With GB200 NVL72 and Kimi-K2.5, MaxLPS reduces the power provisioned per rack in its tests from 125 to 90 kW while maintaining the workload’s throughput. That would allow 39% more racks to be installed within the same power limit.
For Vera Rubin NVL72 running DeepSeek-R1, provisioned power goes from 136 to 101 kW. In this case, NVIDIA calculates that roughly 35% more racks would fit.
The company also puts the performance-per-watt improvement at around 1.5x on GB200 NVL72 and between 1.3x and 1.4x on Vera Rubin NVL72 for the scenarios evaluated. These are NVIDIA’s own results and projections for specific workloads, so they don’t amount to an automatic improvement for any given data center or application.
The company extends the calculation further when it combines MaxLPS with the full electrical and thermal planning of a future Vera Rubin facility.
According to its estimates, a 100 MW data center could house up to 40,000 Rubin GPUs and reach the following capacities:
| Capacity estimated by NVIDIA | 100 MW AI facility |
|---|---|
| NVFP4 inference | 2 zettaFLOPS |
| NVFP4 training | 1.4 zettaFLOPS |
| HBM4 memory | 11 PB |
| HBM4 bandwidth | 800 PB/s |
| Potential GPU increase | Up to 40% |
That 40% should be understood as the maximum target NVIDIA projects by combining several measures, not as an inherent property of every Rubin rack.
Cooling With Warmer Liquid Can Save Electricity
The second part of the proposal sounds contradictory at first: using warmer liquid to cool a data center more effectively.
Vera Rubin NVL72 is designed to run on liquid cooling with an inlet temperature of up to 45°C.
The goal isn’t to cool the chip further, but to reduce the amount of energy needed to later expel that heat from the building.
When the loop runs at higher temperatures, it increases the number of hours during which certain facilities can rely on outside air or free cooling instead of leaning heavily on compressors and chillers.
The outcome depends heavily on climate and on the specific design.
A facility in a cold region will have different opportunities than one built in an environment with high outdoor temperatures for much of the year.
Water use also changes. NVIDIA says a proper design can reduce reliance on water-intensive evaporative cooling, though that doesn’t mean every Rubin data center will do away with chillers or evaporative systems.
The proposal is to use them only when conditions actually require it.
Every kilowatt cooling stops consuming can remain available for the servers, provided the electrical and thermal infrastructure was designed for it.
The Software Side Is Also Chasing More Tokens per Watt
MaxLPS doesn’t stop at the electrical installation.
NVIDIA introduces profiles called Workload Profile Power Solutions (WPPS) to adapt how the GPUs operate to different kinds of workload, including training, inference, memory-bound tasks, or compute-bound jobs.
The Application Performance and Power Manager applies the corresponding settings, and NVIDIA Dynamo can step in to manage distributed inference services.
The goal is to avoid using an identical power configuration for workloads that behave very differently.
For inference, the company proposes measuring results in tokens per second per watt. That metric directly ties consumption to the work performed, though it can’t be used in isolation either: latency, quality of service, and infrastructure cost still matter.
For an AI provider, producing more tokens with the same megawatt can increase available commercial capacity without waiting for a new grid connection.
And that last point is taking on considerable importance.
Electricity Is Starting to Determine the Real Size of AI Clusters
In the early years of generative AI’s growth, the most visible shortage was in GPUs.
Getting hold of accelerators was the problem.
Now the constraints are also shifting toward electrical power, network capacity, transformers, cooling, land, and the time needed to build new facilities.
A company might have the capital needed to buy more GPUs and discover it has nowhere to plug them in.
NVIDIA’s strategy is trying to respond precisely to that situation.
If it takes a data center years to secure another 100 MW, improving the effective use of its existing megawatts by 20%, 30%, or 40% can carry enormous economic value.
It also changes how infrastructure gets measured — a shift that’s already making cost-per-megawatt comparisons less useful on their own.
The number of GPUs still matters, but two data centers with the same power and the same accelerators can produce different amounts of work if one needs more electricity for cooling, keeps more reserved capacity unused, or runs its workloads on a less efficient power configuration.
That’s why NVIDIA is increasingly leaning on the concept of the AI factory: treating the data center as an industrial facility whose output is tokens and other computational results.
It’s marketing language from the manufacturer, but it describes a real shift in how these facilities are being planned.
Designing Now for the GPUs That Will Arrive Later
Another idea within MaxLPS directly affects how new facilities get built.
NVIDIA recommends sizing power, cooling, space, and communications from day one for the maximum number of racks a facility could house over its useful life, even if not all of them get installed initially.
The reason lies in how workloads evolve.
A facility built mainly for training might start with high average power draw per rack. If inference later takes up a larger share and average consumption drops, room would open up to install more hardware without raising the total contracted power.
For that to be possible, the rack positions, electrical distribution, cooling, and network capacity need to physically exist already.
Modifying those elements after a building is already constructed can be far more complex and expensive.
MaxLPS is therefore trying to move the power limit away from being managed GPU by GPU or rack by rack and toward something considered from the data center’s overall design.
The approach still has to prove itself at scale.
Dynamic Power Software and DSX Exchange are currently in Developer Preview, and much of the Vera Rubin figures are NVIDIA’s own projections for future hardware and data centers. They shouldn’t be confused with generalizable results from commercial facilities already in operation.
But the problem it’s trying to solve is real.
The next generation of AI facilities will have GPUs capable of enormous amounts of computation and HBM memory with bandwidth measured in terabytes per second. None of that matters if the facility doesn’t have enough electricity to keep them running.
The next race in AI infrastructure may therefore end up measured less by how many GPUs can be bought and more by how many tokens each megawatt reaching the data center can produce.
Frequently Asked Questions
What is NVIDIA DSX MaxLPS?
DSX MaxLPS is a set of NVIDIA technologies for jointly managing power, GPU configuration, cooling, and data center design. Its goal is to increase the amount of AI work produced within a fixed power limit.
Can MaxLPS install 40% more GPUs without adding electricity?
NVIDIA projects up to 40% more Rubin capacity when MaxLPS is combined with proper data center planning. It’s the manufacturer’s maximum estimate and will depend on the workloads, the facility, and its electrical and thermal design.
Why does Vera Rubin use 45°C liquid cooling?
Running with warmer inlet liquid can make it easier to use outside-air cooling and reduce the hours mechanical cooling systems need to run. The actual savings depend on climate and each facility’s design.
What is Dynamic Power Software?
It’s NVIDIA’s system that monitors consumption and lets available electrical capacity be reallocated among resource groups, racks, nodes, and GPUs within previously set policies and limits. It’s currently in Developer Preview.

