The new MLPerf Inference 6.1 round paints a more complex picture of the race for artificial intelligence infrastructure. AMD has pushed its Instinct MI355X to a 512-GPU configuration that sets a new scale record in the benchmark, while NVIDIA debuts preliminary Vera Rubin results and shows major Blackwell gains. Intel, for its part, expands the presence of Xeon and Arc Pro B70 with gains achieved largely through software.
MLPerf 6.1 in 30 seconds
- AMD reaches 5.75 million tokens/s with 512 MI355X GPUs on GPT-OSS-120B and 2.90 million on DeepSeek-R1.
- NVIDIA debuts Vera Rubin NVL72 with up to 3.7 times the performance of GB300 NVL72 on Qwen3-VL.
- A 288-GPU GB300 system achieves 99% scaling efficiency on DeepSeek-R1.
- In 8-GPU configurations, GB300 outperforms MI355X on several models, though AMD keeps an edge on certain workloads.
- Intel improves Arc Pro B70 and Xeon 6 performance without changing the hardware used in some comparisons.
MLPerf Inference 6.1 also arrives with an important change to the workloads used to measure the machines. MLCommons has added tests aimed at scenarios closer to today’s AI applications, including End-to-End Retrieval-Augmented Generation (E2E-RAG) and Edge Agentic Inference. The first measures a full pipeline that includes embeddings, vector search, reranking, and generation with a language model. The second aims to reflect multi-turn agentic workloads, where each query depends on the previous context.
This round brought together 30 participating organizations, the highest number ever recorded by MLPerf Inference, up from the field that took part in MLPerf Inference v6.0 earlier this year. New processors and accelerators also appear, including the AMD Ryzen AI Max+ 395, Instinct MI350P, Intel Arc Pro B70, and NVIDIA’s first Rubin platforms, the latter still in preview mode.
AMD scales the MI355X to 512 GPUs
The most striking figure among the results analyzed is the 512 AMD Instinct MI355X configuration submitted by Crusoe. On GPT-OSS-120B it reaches 5.75 million tokens per second in the Offline scenario and 5.39 million in Server. On DeepSeek-R1 it hits 2.90 million tokens per second in Offline and 2.41 million in Server. AMD describes these results as the highest aggregate token throughput recorded so far in MLPerf for those workloads.
The comparison calls for some caution, because the number of accelerators isn’t equal across all configurations. On DeepSeek-R1, for example, the 512-GPU AMD configuration reaches 2,901,950 tokens/s in the published data, while a 288-GPU NVIDIA GB300 configuration posts 2,705,130 tokens/s. In this case, AMD achieves higher aggregate throughput with more accelerators, so it can’t be read as a direct per-GPU efficiency comparison.
The picture changes when the scale shrinks.
On GPT-OSS-120B, the eight-GPU MI355X reaches 117,804 tokens/s in Offline and 110,188 in Server according to the data in the charts. The eight-GPU GB300 hits 132,236 and 127,921 tokens/s, respectively. In other words, NVIDIA’s system keeps an edge in that specific configuration.
AMD also shows a considerable improvement as the number of MI355X GPUs grows. The company reports 95% scaling efficiency between eight and 72 GPUs on GPT-OSS-120B, a figure indicating that a large share of the theoretical additional performance is retained when moving to a multi-node configuration.
That behavior matters more for large infrastructures than an isolated peak figure. As a platform grows from a single server to multiple racks, GPU-to-GPU communication, memory, request distribution, and software can all become bottlenecks.
Vera Rubin makes its first MLPerf appearance
The round’s other major development comes from NVIDIA. Vera Rubin NVL72 appears in MLPerf Inference for the first time through preview results — a debut covered in more detail in our full breakdown of the Vera Rubin NVL72 numbers against GB300.
On DeepSeek-R1, the 72-GPU VR200 system posts 1,183,327 tokens/s in Offline, compared to 689,961 tokens/s for a 72-GPU GB300 NVL72 system. On Qwen3-VL, Vera Rubin reaches 1,307 queries per second versus 349 for GB300, roughly 3.7 times more. NVIDIA’s documentation flags the Vera Rubin results as preliminary.
The comparison with Blackwell also shows that scale matters. A 72-GPU GB300 system reaches 679,740 tokens/s on DeepSeek-R1, while a 288-GPU configuration hits 2,705,130. NVIDIA says the jump from one to four NVL72 platforms keeps 99% scaling efficiency in the Offline scenario.
That 99% doesn’t mean four racks are four times faster on any given workload. It refers to the scaling efficiency of a specific configuration within the DeepSeek-R1 benchmark — a metric that measures how much additional performance is gained as available resources increase.
NVIDIA also credits part of the improvement between MLPerf 6.0 and 6.1 to its software. The company estimates gains of up to 1.6x on certain results thanks to changes in kernels, operation fusion, KV-cache precision, and distributed serving techniques.
Intel shows software can move the needle too
Intel shows up in this round on two separate fronts: Xeon 6 processors and Arc Pro B70 GPUs.
On Xeon 6, the company expands from two to five evaluated models compared with MLPerf 6.0. Intel says a Xeon 6980P achieved 2.4x the Server performance on Llama 3.1 8B compared with the previous round using the same silicon and socket count, while Offline performance grew 56%.
With Arc Pro B70, a machine equipped with four GPUs packs 128 GB of VRAM and appears in tests with Llama 3.1 8B, Llama 2 70B, GPT-OSS-120B, Whisper, and E2E-RAG.
On GPT-OSS-120B, Intel compares the same four-Arc-Pro-B70 system used in MLPerf 6.0 and reports a 36% improvement in Server and 27% in Offline. Results climb from 951.67 to 1,296.87 tokens/s in Server and from 1,536.90 to 1,956.40 in Offline.
The data in the charts also place the Arc Pro B70 below AMD’s and NVIDIA’s larger GPUs on the LLM workloads analyzed. On Llama 3.1 8B, for example, four Arc Pro B70 GPUs reach 4,977 tokens/s in Offline, compared with 158,458 for the eight-GPU MI355X and 171,114 for the eight-GPU GB300.
The comparison isn’t meant to place these machines in the same category. In fact, one of the more interesting aspects of MLPerf 6.1 is that the benchmark is starting to capture configurations spanning everything from large multi-node systems to smaller workstations and servers.
The numbers change when the model changes
MLPerf 6.1’s results also show why a single figure isn’t enough to talk about “the fastest GPU.”
On Llama 2 70B, an eight-GPU GB300 system reaches 132,253 tokens/s in Offline versus 103,792 for the eight-GPU MI355X. On Llama 3.1 8B, GB300 hits 171,114 versus 158,458 for MI355X. But on GPT-OSS-120B, the relationships shift depending on the configuration and scenario.
The result also changes as the number of accelerators grows. On DeepSeek-R1, the 512-GPU AMD system beats the 288-GPU NVIDIA configuration mentioned above in aggregate throughput, while Vera Rubin posts far higher results than GB300 on certain tests using only 72 GPUs. These are useful comparisons for understanding how each platform behaves, but they don’t add up to a universal ranking.
There’s another factor complicating the picture even further: software.
MLPerf compares specific configurations, software versions, and given models. NVIDIA, AMD, and Intel are continuously optimizing their inference stacks. That’s why a result published today can be surpassed by a different version of kernels, compilers, inference libraries, or serving systems without the GPU itself changing at all.
MLCommons is pointing precisely in that direction with the benchmark’s evolution. More than half of this round’s participants used the new API-based test harness, designed as the foundation for the future MLPerf Endpoints family and built around a client-server architecture closer to certain real-world data center deployments.
The 6.1 round leaves behind a race split across several fronts. AMD shows the MI355X can scale to 512 GPUs with a token volume that sets new aggregate records; NVIDIA uses Vera Rubin to give a first look at its next generation within MLPerf while keeping Blackwell competitive; and Intel shows there’s still room to extract more performance from Xeon and Arc Pro through software.
For those designing AI infrastructure, the peak tokens-per-second figure is only part of the equation. The model used, system size, latency, scaling efficiency, precision, memory, software, and the cost of operating the whole system determine which result actually matters for a given workload. MLPerf 6.1 leaves exactly that idea on the table: the competition is no longer just about having the fastest GPU, but about keeping the entire system’s performance intact as inference moves from the lab to hundreds of accelerators.

