PulseAugur
实时 07:35:20
实体 Nemotron-Labs-3-Puzzle-75B-A9B

Nemotron-Labs-3-Puzzle-75B-A9B

PulseAugur coverage of Nemotron-Labs-3-Puzzle-75B-A9B — every cluster mentioning Nemotron-Labs-3-Puzzle-75B-A9B across labs, papers, and developer communities, ranked by signal.

Show in brief
总计 · 30天
0
90 天内 4
发布 · 30天
0
90 天内 0
论文 · 30天
0
90 天内 0
层级分布 · 90 天
主题
时间线
  1. 2026-07-07 research_milestone Researchers released Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant of Nemotron-3-Super optimized for interactive deployment and long-context tasks. 来源
  2. 2026-07-07 research_milestone NVIDIA researchers released Nemotron-Labs-3-Puzzle-75B-A9B, a compressed LLM variant demonstrating significant improvements in inference efficiency and concurrency. 来源
最近 · 第 1/1 页 · 共 4 条
  1. RESEARCH · CL_145980 ·

    NVIDIA发布新的嵌入模型,用户演示在消费级硬件上运行大模型

    NVIDIA发布了两个新的文本嵌入模型Nemotron-3-Embed-8B-BF16和Nemotron-3-Embed-1B-BF16,它们针对检索和语义相似性任务进行了优化。这些模型专为多语言应用和检索增强生成(RAG)系统设计,在相关基准测试中取得了最先进的性能。此外,一位用户已成功在消费级硬件上部署了Nemotron-Labs-3-Puzzle-75B-A9B模型,展示了其在大上下文窗口和高效推理方面的能力。

  2. FRONTIER RELEASE · CL_130046 ·

    NVIDIA 发布 Nemotron-Labs-3-Puzzle-75B 以支持 Blackwell 硬件

    NVIDIA 发布了其 Nemotron-Labs-3-Puzzle-75B 模型,该模型已针对 Blackwell 硬件上的服务进行了优化。该模型集成了 LatentMoE、Mamba-Interleaving 和 Multi-Token Prediction (MTP) 以提高吞吐量。它在 OpenMDW-1.1 许可下可用,允许商业使用。

  3. RESEARCH · CL_128708 ·

    NVIDIA压缩Nemotron-3大语言模型,吞吐量提升2倍,100万token并发提升8倍

    NVIDIA研究人员开发了Nemotron-Labs-3-Puzzle-75B-A9B,这是其Nemotron-3-Super大语言模型的压缩版本。该新变体显著提高了部署效率,在8xB200节点上实现了高达2倍的服务器吞吐量,并在单个H100 GPU上实现了高达8个并发100万token请求。通过结合迭代式Puzzle压缩、知识蒸馏和量化等技术的阶段性流水线实现了压缩,同时在很大程度上保留了模型在下游任务中的准确性。

  4. FRONTIER RELEASE · CL_129798 ·

    NVIDIA发布压缩版Nemotron大语言模型,增强音频和推理能力 · 跟踪10个来源

    NVIDIA发布了基于其Nemotron架构的几款新模型,包括统一的音频-文本大语言模型Nemotron-Labs-Audex-2B和Nemotron-Labs-Audex-30B-A3B。此外,Nemotron-Labs-3-Puzzle-75B-A9B是Nemotron-3-Super-120B-A12B的部署优化压缩版本,旨在提高推理效率和吞吐量。这些模型正与LangChain和Fireworks等框架集成,以实现高级代理能力,…