NVIDIA has released new early results for Vera Rubin NVL72 that point to a big leap in inference efficiency for agentic AI. In tests it ran with SemiAnalysis’s AgentX benchmark, the new platform reaches up to 30 times more throughput per megawatt than GB300 NVL72, and a cost per million tokens up to 35 times lower under certain conditions. The results still await review by SemiAnalysis.
NVIDIA Vera Rubin in 30 seconds
- Vera Rubin NVL72 reaches up to 30 times higher throughput per megawatt than GB300 NVL72 on DeepSeek V4 Pro.
- NVIDIA estimates up to 35 times lower cost per million tokens than Blackwell in these tests.
- AgentX simulates real programming sessions with agents, with growing context and tool calls.
- Blackwell improves too: GB300 reaches up to 15 times the performance per MW of H200 on DeepSeek V4 Pro.
- Vera Rubin’s figures were measured by NVIDIA and still await validation by SemiAnalysis.
The comparison is interesting because NVIDIA isn’t just measuring how many tokens a GPU produces in a static test. AgentX tries to mirror how new AI agents actually work, where one task can involve reasoning, tool calls, sub-agents, and many requests with context that grows through the session.
It also shifts which metrics matter in data centers. Peak accelerator performance still counts, but for operators running dozens or hundreds of racks, tokens per megawatt, latency, time to first token, and cost per million tokens matter more and more.
Vera Rubin reaches up to 30 times the throughput per megawatt
AgentX is part of InferenceX, SemiAnalysis’s open benchmark suite. Instead of only using artificial fixed-length sequences, it replays recorded programming sessions with agents.
The tests keep elements like context growth, variable input and output lengths, reasoning times, and pauses from tool calls. They also vary concurrency to show the relationship between overall performance and responsiveness per user.
On DeepSeek V4 Pro, NVIDIA says Vera Rubin NVL72 reaches up to 30 times more throughput per megawatt than GB300 NVL72 when comparing at around 160 tokens per second per user.
That doesn’t mean Vera Rubin is 30 times faster than Blackwell across all workloads. The figure sits at a specific point on the AgentX curve, with a particular mix of model, concurrency, and interactivity.
The distinction matters because NVIDIA earlier announced gains of about 10 times in performance per watt and cost per token for Vera Rubin over Blackwell in other inference scenarios. The new results don’t contradict those; they come from a different AI load and method.
In AgentX, the platform can also reach higher interactivity than GB300 NVL72. NVIDIA’s published charts show Vera Rubin still scaling into ranges where Blackwell starts to hit its limits.
The company credits a broader system design, not just the Rubin GPU. The Vera Rubin architecture is built as an integrated system: GPU, CPU, memory, interconnects, networks, and software work together to cut data movement, waiting, and recomputation.
Up to 35 times lower cost per million tokens
The other headline number has more direct economic weight. NVIDIA says Vera Rubin NVL72 can reach up to 35 times lower cost per million tokens than GB300 NVL72 in the new AI load tests.
Again, these are maximums under specific conditions NVIDIA used, not a blanket cut for any AI service moving from Blackwell to Rubin.
But the metric shows where AI infrastructure is heading.
In the early phase of generative AI, much of the focus was on buying as many GPUs as possible. With large-scale inference, the question is more about how much useful work you get per GPU, and especially per available megawatt in the data center.
For operators with limited electrical capacity, getting more tokens from the same energy budget means handling more requests without scaling electrical and cooling infrastructure in proportion.
NVIDIA even extends this to data center energy management. Its DSX MaxLPS technology manages power at the GPU, rack, and workload level, and reportedly can let operators install up to 40% more GPUs within the same megawatt budget.
That makes available electricity one of the key measures of what an AI infrastructure can really do.
Blackwell also improves over Hopper
The new data isn’t only about Rubin. NVIDIA used AgentX to compare its current Blackwell generation with Hopper.
With DeepSeek V4 Pro 1.6T, GB300 NVL72 reaches up to 15 times more throughput per megawatt than H200 NVL8, per the published results. Cost per million tokens can be ten times lower.
The gap grows with bigger models.
In tests with Kimi K3 2.8T, NVIDIA puts GB300 NVL72 around 80 times higher in throughput per megawatt than H200 NVL8 at comparable interactivity. Blackwell can also push interactive performance to roughly 215 tokens per second per user in this scenario.
Part of the gains come from hardware, but NVIDIA also credits software. Tools like SGLang, TensorRT-LLM, and vLLM help distribute Mixture of Experts (MoE) models, along with DeepGEMM kernels, reduced-precision formats like MXFP4 and MXFP8, and NVIDIA Dynamo.
Dynamo can split the prefill and decode phases across different resource groups, and its routing takes existing KV caches into account to avoid reprocessing parts of the context.
NVLink ties the architecture together, connecting the 72 GPUs in an NVL72 as a high-speed compute domain. The Rubin generation raises NVLink bandwidth in NVL72 from Blackwell’s 130 TB/s to 260 TB/s, per NVIDIA’s preliminary specs.
The Vera Rubin results out now don’t yet show the platform’s full potential. NVIDIA notes these measurements don’t include Vera CPU performance for tool calls, which matters especially for agents that alternate token generation with running code or using external services.
So the benchmark snapshot is promising but still preliminary. The figures of up to 30 times more throughput per megawatt and 35 times lower cost come from NVIDIA’s own AgentX measurements and await independent review by SemiAnalysis. That validation, plus real deployments, will show how well these advantages carry into production.
Frequently Asked Questions
How much does NVIDIA Vera Rubin improve over Blackwell?
NVIDIA reports up to 30 times higher throughput per megawatt with Vera Rubin NVL72 than GB300 NVL72 on DeepSeek V4 Pro in AgentX. That’s a maximum under specific benchmark conditions, not a universal 30x gain for every application.
Does Vera Rubin really cut token costs by 35 times?
NVIDIA claims up to 35 times lower cost per million tokens with Vera Rubin in the new AI load tests. The results still await review by SemiAnalysis.
What is AgentX?
AgentX is an open benchmark from SemiAnalysis, part of InferenceX, that replays programming sessions with agents. It accounts for growing context, tool calls, varying input and output lengths, KV caching, concurrency, and latency.
Why does throughput per megawatt matter in AI?
Data centers have limited electrical power. Producing more tokens per megawatt lets you scale inference capacity without raising energy use in proportion.
via: blogs.nvidia

