PulseAugur
EN
LIVE 10:31:14

NVIDIA compresses Nemotron-3 LLM for 2x throughput, 8x 1M-token concurrency

NVIDIA researchers have developed Nemotron-Labs-3-Puzzle-75B-A9B, a compressed version of their Nemotron-3-Super large language model. This new variant significantly enhances deployment efficiency, achieving up to twice the server throughput on 8xB200 nodes and enabling up to 8 concurrent 1M-token requests on a single H100 GPU. The compression was achieved through a multi-stage pipeline combining techniques like iterative puzzle compression, knowledge distillation, and quantization, while largely preserving the model's downstream accuracy. AI

IMPACT This compression technique could enable more efficient deployment of large language models, increasing throughput and concurrency for interactive applications.

RANK_REASON Research paper detailing a new compressed LLM variant with performance improvements.

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

NVIDIA compresses Nemotron-3 LLM for 2x throughput, 8x 1M-token concurrency

COVERAGE [2]

  1. arXiv cs.AI TIER_1 English(EN) · Akhiad Bercovich, Talor Abramovich, Daniel Afrimi, Shay Aharon, Nir Ailon, Vladimir Anisimov, Omer Ullman Argov, Maor Ashkenazi, Tomer Asida, Nave Assaf, Tomer Bar Natan, Alexander Bukharin, Grzegorz Chlebus, Marcin Chochowski, Eric Chung, Mohammad Dabba… ·

    Nemotron-Labs-3-Puzzle-75B-A9B: Compressing Hybrid MoE LLMs

    arXiv:2607.04371v1 Announce Type: new Abstract: We present Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant of Nemotron-3-Super optimized for interactive deployment. We designed the model to maximize server throughput under high user throughput constraints. In interactive ser…

  2. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    Meet Nemotron Labs 3 Puzzle 75B A9B: A Compressed Hybrid MoE LLM Delivering 2.03x Server Throughput

    <p>NVIDIA has released Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant of Nemotron-3-Super. Iterative Puzzle alternates hardware-aware structural compression with short knowledge distillation recovery phases. The model drops from 120.7B total / 12.8B active parameters to 75.…