NVIDIA has announced leading performance for its new Vera Rubin NVL72 system in the MLPerf Inference v6.1 benchmarks. The system demonstrated up to 3.7x higher throughput than its predecessor, the GB300 NVL72, on demanding models like Qwen3-VL and DeepSeek-R1. This performance is attributed to full-stack codesign across hardware and software, including enhanced Tensor Cores, Transformer Engine, and optimized interconnects. NVIDIA also highlighted the system's efficient scaling capabilities, achieving 99% efficiency in a multi-rack submission, and its potential to significantly improve AI inference economics by generating more tokens and serving more users at a lower cost. AI
IMPACT Sets new performance benchmarks for AI inference hardware, potentially lowering costs and increasing efficiency for AI deployments.
RANK_REASON Debut submission of a new hardware system to a recognized industry benchmark. [lever_c_demoted from research: ic=1 ai=0.7]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →