NVIDIA has released new early results for Vera Rubin NVL72 indicating a significant leap in inference efficiency for intelligent AI. In tests conducted by the company using SemiAnalysis’s AgentX benchmark, the new platform achieves up to 30 times more throughput per megawatt than GB300 NVL72 and a cost per million tokens that’s up to 35 times lower under certain conditions. These results are still pending review by SemiAnalysis.
The key highlights of NVIDIA Vera Rubin in 30 seconds
- Vera Rubin NVL72 reaches up to 30 times higher throughput per megawatt than GB300 NVL72 in DeepSeek V4 Pro.
- NVIDIA estimates a cost reduction of up to 35 times per million tokens compared to Blackwell in these tests.
- AgentX simulates real programming sessions with agents, increasing context and tool calls.
- Blackwell also sees improvements: GB300 reaches up to 15 times the performance per MW of H200 with DeepSeek V4 Pro.
- Vera Rubin’s figures were measured by NVIDIA and are still awaiting validation by SemiAnalysis.
This comparison is especially interesting because NVIDIA is not just measuring how many tokens a GPU can generate in a static test. AgentX aims to replicate a scenario much closer to how new AI agents operate, where a single task can involve reasoning, tool calls, sub-agents, and multiple requests with an increasing context during the session.
This also shifts the metrics that start to matter in data centers. While maximum accelerator performance remains relevant, for operators deploying dozens or hundreds of racks, tokens per megawatt, latency, time to first token, and cost per million tokens are increasingly crucial.
Vera Rubin achieves up to 30 times the throughput per megawatt
AgentX is part of InferenceX, SemiAnalysis’s open benchmark suite. Instead of only using artificial sequences of fixed length, it reproduces recorded programming sessions involving agents.
The tests retain elements like context growth, variable input/output lengths, reasoning times, and pauses caused by tool calls. They also manipulate concurrency to observe the relationship between overall performance and responsiveness for each user.
In DeepSeek V4 Pro, NVIDIA claims Vera Rubin NVL72 achieves up to 30 times more throughput per megawatt than GB300 NVL72 when comparing around 160 tokens per second per user.
This does not mean Vera Rubin is 30 times faster than Blackwell for all workloads. That figure corresponds to a specific point on the AgentX performance curve, involving a particular combination of model, concurrency, and interactivity.
This precision is important because NVIDIA previously announced improvements of about 10 times in performance per watt and cost per token for Vera Rubin compared to Blackwell in other inference scenarios. The new results do not contradict those figures; they stem from a different AI load and methodology.
In AgentX, the platform can also reach higher levels of interactivity than GB300 NVL72. NVIDIA’s published graphs show Vera Rubin continuing to scale into ranges where Blackwell begins to reach its limits.
The company attributes this to a broader system design beyond just the Rubin GPU. The Vera Rubin architecture is built around an integrated system: GPU, CPU, memory, interconnects, networks, and software work together to reduce data movement, waiting, and recomputation.
Up to 35 times less cost per million tokens
The other highlighted figure has more direct economic implications. NVIDIA states that Vera Rubin NVL72 can achieve up to 35 times lower cost per million tokens than GB300 NVL72 in the new AI load tests.
Again, these are maximums based on specific conditions used by NVIDIA, not a blanket reduction for any AI service migrating from Blackwell to Rubin.
But the metric helps illustrate where AI infrastructure is heading.
During the early phase of generative AI, much focus was on acquiring as many GPUs as possible. With large-scale inference, the equation is increasingly about how much useful work can be extracted per GPU, especially per available megawatt in the data center.
For operators with limited electrical capacity, multiplying tokens produced within the same energy budget allows more requests to be handled without proportionally increasing electrical and cooling infrastructure.
NVIDIA is even extending this approach to data center energy management. Its DSX MaxLPS technology manages power at the GPU, rack, and workload level, and reportedly can enable installing up to 40% more GPUs within the same megawatt budget.
This makes available electricity one of the key metrics to measure the true capacity of an AI infrastructure.
Blackwell also improves over Hopper
The new data isn’t limited to Rubin. NVIDIA has used AgentX to measure how its current Blackwell generation compares to Hopper.
With DeepSeek V4 Pro 1.6T, GB300 NVL72 achieves up to 15 times more throughput per megawatt than H200 NVL8, according to published results. Cost per million tokens may be ten times lower.
The difference grows with larger models.
In tests with Kimi K3 2.8T, NVIDIA places GB300 NVL72 around 80 times higher in throughput per megawatt than H200 NVL8 at comparable levels of interactivity. Blackwell can also extend interactive performance to roughly 215 tokens per second per user in this scenario.
Part of these improvements comes from hardware, but NVIDIA credits software tools as well. Technologies like SGLang, TensorRT-LLM, and vLLM facilitate distribution of Mixture of Experts (MoE) models, along with DeepGEMM kernels, reduced-precision formats like MXFP4 and MXFP8, and NVIDIA Dynamo.
Dynamo can divide the prefill and decode phases across different resource groups, and its routing system considers existing KV caches to avoid reprocessing parts of the context.
NVLink completes this architecture by connecting the 72 GPUs in an NVL72 as a high-speed computing domain. The Rubin generation increases NVLink bandwidth in NVL72 from the 130 TB/s of Blackwell to 260 TB/s, per NVIDIA’s preliminary specs.
The Vera Rubin results released now do not yet showcase all of the platform’s potential. NVIDIA notes these measurements do not include Vera CPU performance for tool calls, which is especially relevant for agents that alternate token generation with code execution or external service use.
Thus, the benchmark snapshot is promising but still preliminary. The figures of up to 30 times more throughput per megawatt and 35 times lower cost come from NVIDIA measurements on AgentX and await independent review by SemiAnalysis. That validation, along with commercial deployments, will determine how well these advantages transfer to production infrastructures.
Frequently Asked Questions
How much does NVIDIA Vera Rubin improve over Blackwell?
NVIDIA reports up to 30 times higher throughput per megawatt with Vera Rubin NVL72 compared to GB300 NVL72 using DeepSeek V4 Pro in AgentX. This figure represents a maximum under specific benchmark conditions, not a universal 30x improvement for all applications.
Does Vera Rubin truly reduce token costs by 35 times?
NVIDIA claims to have achieved up to 35 times lower cost per million tokens with Vera Rubin in the new AI load tests. These results are still pending review by SemiAnalysis.
What is AgentX?
AgentX is an open benchmark from SemiAnalysis integrated into InferenceX that reproduces programming sessions with agents. It considers increasing context, tool calls, varying input/output lengths, KV caching, concurrency, and latency.
Why does throughput per megawatt matter in AI?
Data centers have limited electrical power. Producing more tokens per megawatt allows scaling inference capacity without proportionally increasing energy consumption.
via: blogs.nvidia

