Vera Rubin NVL72 debuts in MLPerf Inference v6.1 at up to 3.7x GB300 on Qwen3-VL
NVIDIA Blog·medium signal
NVIDIA posted Vera Rubin NVL72's first MLPerf Inference results on 16 September across DeepSeek-R1, Qwen3-VL, WAN 2.2 text-to-video, GPT-OSS-120B, DLRMv3 and an edge-agentic Qwen3.6-27B test. Vera Rubin NVL72 delivered up to 3.7x the throughput of GB300 NVL72 on Qwen3-VL and up to 2.5x on DeepSeek-R1. Separately, GB300 NVL72 scaled from 72 to 288 GPUs at 99% efficiency on DeepSeek-R1 offline, and software alone (kernel fusion, disaggregated serving) moved GB300's Qwen3-VL numbers up to 1.6x between v6.0 and v6.1.