NVIDIA's Rubin Platform Sets New Inference Benchmark in MLPerf v6.1

In its first MLPerf Inference v6.1 submission, NVIDIA's Vera Rubin NVL72 system achieved up to 3.7x higher throughput than the GB300 NVL72 on the Qwen3-VL benchmark and 2.5x on DeepSeek-R1. The company also reported 99% scaling efficiency across four GB300 NVL72 racks, with software optimizations contributing up to 1.6x performance gains over the previous version.
The Vera Rubin NVL72's debut marks the first preview submission for this next-generation platform, with results drawn from the Closed Division of the MLPerf suite. Performance gains stem from full-stack engineering: enhanced Tensor Cores and Transformer Engine speed both prefill and decode phases, while NVFP4 precision shrinks memory demands across weights, attention layers, and KV cache. Disaggregated serving—separating prefill from decode—combined with large-scale expert parallelism proved critical for mixture-of-experts models.
Scaling efficiency also featured prominently. Four GB300 NVL72 racks operating as a 288-GPU cluster achieved 99% scaling efficiency, meaning throughput grew nearly linearly from a single-rack baseline. NVIDIA additionally credited continuous software refinement, with optimizations in v6.1 delivering up to 1.6x performance improvements over the prior version, and further gains emerging even after submission. The NVL72's sixth-generation NVLink and NVLink Switch provide 10x higher packet rates and 3x lower latency than standard Ethernet.
These results could reshape how enterprises evaluate AI infrastructure investments, as higher throughput per rack translates directly into lower cost per token and faster return on hardware spending. Organizations running large-scale inference workloads—cloud providers, AI startups, and enterprise deployments—may find themselves reassessing upgrade cycles, potentially accelerating adoption of next-generation platforms. Meanwhile, the emphasis on software-driven gains suggests existing hardware owners could extend system lifespans through optimization alone, softening the pressure to purchase new equipment. The broader effect may be intensified competition among chip vendors to demonstrate measurable inference efficiency rather than raw specifications.