PulseAugur
实时 08:26:27
English(EN) Nemotron-Labs-3-Puzzle-75B-A9B: Compressing Hybrid MoE LLMs

NVIDIA压缩Nemotron-3大语言模型,吞吐量提升2倍,100万token并发提升8倍

NVIDIA研究人员开发了Nemotron-Labs-3-Puzzle-75B-A9B,这是其Nemotron-3-Super大语言模型的压缩版本。该新变体显著提高了部署效率,在8xB200节点上实现了高达2倍的服务器吞吐量,并在单个H100 GPU上实现了高达8个并发100万token请求。通过结合迭代式Puzzle压缩、知识蒸馏和量化等技术的阶段性流水线实现了压缩,同时在很大程度上保留了模型在下游任务中的准确性。 AI

影响 这项压缩技术可以实现更高效的大语言模型部署,提高交互式应用的吞吐量和并发性。

排序理由 研究论文详细介绍了一种新的压缩大语言模型变体,并带来了性能提升。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

NVIDIA压缩Nemotron-3大语言模型,吞吐量提升2倍,100万token并发提升8倍

报道来源 [2]

  1. arXiv cs.AI TIER_1 English(EN) · Akhiad Bercovich, Talor Abramovich, Daniel Afrimi, Shay Aharon, Nir Ailon, Vladimir Anisimov, Omer Ullman Argov, Maor Ashkenazi, Tomer Asida, Nave Assaf, Tomer Bar Natan, Alexander Bukharin, Grzegorz Chlebus, Marcin Chochowski, Eric Chung, Mohammad Dabba… ·

    Nemotron-Labs-3-Puzzle-75B-A9B:压缩混合MoE大语言模型

    arXiv:2607.04371v1 Announce Type: new Abstract: We present Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant of Nemotron-3-Super optimized for interactive deployment. We designed the model to maximize server throughput under high user throughput constraints. In interactive ser…

  2. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    认识 Nemotron Labs 3 Puzzle 75B A9B:一款压缩混合 MoE LLM,服务器吞吐量提升 2.03 倍

    <p>NVIDIA has released Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant of Nemotron-3-Super. Iterative Puzzle alternates hardware-aware structural compression with short knowledge distillation recovery phases. The model drops from 120.7B total / 12.8B active parameters to 75.…