NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut
System performance, efficient infrastructure scaling and continuous software optimization are key levers that determine AI inference economics. Higher system performance means more tokens generated, resulting in higher revenue. Efficient scaling means throughput grows proportionally as hardware gets added, requiring fewer resources to serve users at scale. Continuous optimization means generating…
NVIDIA's latest AI inference system, NVIDIA Vera Rubin NVL72, has made a significant debut in the MLPerf Inference v6.1 competition, delivering leading performance in terms of throughput, scaling efficiency, and software velocity. This achievement highlights NVIDIA's commitment to optimizing across the full stack of AI infrastructure, from hardware to software.
Key factors contributing to the system's success include higher system performance that translates to more tokens generated, leading to higher revenue. The efficient scaling of throughput grows proportionally with added hardware, enabling fewer resources to serve users at scale. Continuous software optimizations have further enhanced the performance, with up to 1.6x higher performance over the v6.0 version.
On the benchmark tests, NVIDIA Vera Rubin NVL72 outperformed the NVIDIA GB300 NVL72 by up to 3.7x on Qwen3-VL and 2.5x on DeepSeek-R1, showcasing the system's capability to serve more users and generate more revenue, while reducing cost per token. These improvements are attributed to NVIDIA's enhanced Tensor Cores and Transformer Engine, which accelerate the prefill and decode stages of inference, and NVFP4 precision that reduces memory footprint across model weights, attention, and KV cache.
Additionally, NVIDIA's disaggregated serving technique, which separates prefill and decode along with large-scale expert parallelism, further optimizes the system's performance across mixture-of-experts layers in models like DeepSeek-R1 and Qwen3-VL. The NVL72 scale-up domain, powered by NVIDIA's sixth-generation NVLink and NVLink Switch, provides a high-performance interconnect foundation, delivering 10x higher packet rates and 3x lower latency than off-the-shelf Ethernet.
The system's scaling efficiency is a critical measure of AI infrastructure productivity, and NVIDIA has achieved 99% scaling efficiency with the GB300 NVL72, achieving nearly linear throughput growth from a single-rack baseline as more GPUs are added. This efficiency is crucial, as more GPUs do not automatically translate to proportionally more throughput. The architecture, interconnect, and software must all scale together to achieve optimal results.
Overall, NVIDIA Vera Rubin NVL72's debut in MLPerf Inference v6.1 with leading performance in throughput, scaling efficiency, and software velocity underscores NVIDIA's commitment to continuous optimization and innovation in AI infrastructure.
Written by urgent.news from NVIDIA Blog's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.