PulseAugur
实时 06:37:48

NVIDIA发布压缩版Nemotron大语言模型,增强音频和推理能力 · 跟踪10个来源

NVIDIA发布了基于其Nemotron架构的几款新模型,包括统一的音频-文本大语言模型Nemotron-Labs-Audex-2B和Nemotron-Labs-Audex-30B-A3B。此外,Nemotron-Labs-3-Puzzle-75B-A9B是Nemotron-3-Super-120B-A12B的部署优化压缩版本,旨在提高推理效率和吞吐量。这些模型正与LangChain和Fireworks等框架集成,以实现高级代理能力,提供显著的性能提升和成本降低。 AI

影响 这些发布增强了音频处理能力,并显著提高了推理效率,可能降低部署高级大语言模型的成本。

排序理由 来自前沿实验室(NVIDIA)的多款新模型发布,包括音频-文本能力和用于提高效率的压缩版本。

在 Hugging Face Trending Models 阅读 →

AI 生成摘要 · Google Gemini · 来自 20 个来源。 我们如何撰写摘要 →

NVIDIA发布压缩版Nemotron大语言模型,增强音频和推理能力 · 跟踪10个来源

报道来源 [20]

  1. Hugging Face Blog TIER_1 English(EN) ·

    NVIDIA Nemotron 3 Embed 在 RTEB 上综合排名第一,推动 Agentic Retrieval

  2. Hugging Face Trending Models TIER_1 (CA) · nvidia ·

    nvidia/Nemotron-Labs-Audex-2B

    text-generation · 1,050 downloads · 50 likes

  3. Hugging Face Trending Models TIER_1 English(EN) · nvidia ·

    nvidia/Nemotron-Labs-Audex-30B-A3B

    text-generation · 0 downloads · 46 likes

  4. Hugging Face Trending Models TIER_1 Italiano(IT) · nvidia ·

    nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4

    text-generation · 47 downloads · 53 likes

  5. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    了解为 NVIDIA Nemotron 3 Ultra 优化的 LangChain Deep Agents Harness:https://t.co/w1aOqdBazc

    Learn more about the LangChain Deep Agents Harness optimized for NVIDIA Nemotron 3 Ultra: https://t.co/w1aOqdBazc

  6. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    我们的新食谱通过 @LangChain Deep Agents 使用 @nvidia Nemotron 3 Ultra 实现 Ralph 循环:每次迭代都使用文件系统作为内存来获取新上下文

    Our new cookbook implements Ralph loops with @nvidia Nemotron 3 Ultra via @LangChain Deep Agents: fresh context every iteration using the filesystem as memory between iterations. Full notebook: https://t.co/j6BR187vN5 https://t.co/0ZdOyPA50I

  7. X — Fireworks (inference infra) TIER_1 English(EN) · FireworksAI_HQ ·

    ICYMI:@LangChain Deep Agents 在 @nvidia Nemotron 3 Ultra 上运行。前沿开放模型代理的成本比闭源模型低约 10 倍。

    ICYMI: @LangChain Deep Agents on @nvidia Nemotron 3 Ultra. frontier open-model agents at ~10x lower cost than closed. Run on Fireworks, then post-train it into specialized intelligence you own. https://t.co/IiMMr94QLa https://t.co/RZoQp40iTk

  8. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    NVIDIA AI 发布 Nemotron 3 Embed:一个在 RTEB 上排名第一的 8B 检查点开放嵌入集合

    <p>NVIDIA released Nemotron 3 Embed on July 15 and 16, 2026. The collection has three open checkpoints: Nemotron-3-Embed-8B-BF16, Nemotron-3-Embed-1B-BF16, and Nemotron-3-Embed-1B-NVFP4. The 8B ranks #1 on RTEB at 78.46 average NDCG@10. The 1B came from ModelOpt NAS pruning plus …

  9. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    NVIDIA 发布 Nemotron-Labs-3-Puzzle-75B-A9B:压缩混合 MoE LLM 在匹配用户吞吐量下实现 2.03 倍服务器吞吐量

    <p>NVIDIA has released Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant of Nemotron-3-Super. Iterative Puzzle alternates hardware-aware structural compression with short knowledge distillation recovery phases. The model drops from 120.7B total / 12.8B active parameters to 75.…

  10. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    NVIDIA发布Nemotron 3 Embed,一个开源嵌入模型集合。80亿参数版本在检索任务的RTEB基准测试中名列前茅,

    NVIDIA has unveiled Nemotron 3 Embed, an open-source embedding model collection. The 8 billion parameter version tops the RTEB benchmark for retrieval tasks, marking a significant advancement in AI infrastructure for enterprise search and RAG applications. https://www. marktechpo…

  11. dev.to — LLM tag TIER_1 English(EN) · Pneumetron ·

    NVIDIA 发布 Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4:一款为部署优化的混合 MoE LLM

    <h2> What Changed </h2> <p>NVIDIA has introduced Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4, a new large language model (LLM) designed for enhanced inference efficiency. This model is a compressed variant of the Nemotron-3-Super-120B-A12B, specifically optimized for deployment in inter…

  12. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    🧠 新模型:Nemotron 3 Ultra (NVIDIA) 上下文:100万 token · $0.5 进 / $2.2 出 每百万 token 最新模型,每小时追踪:https:// opensourceai.tech/latest.ht

    🧠 New model: Nemotron 3 Ultra (NVIDIA) Context: 1M tokens · $0.5 in / $2.2 out per M All the latest models, tracked hourly: https:// opensourceai.tech/latest.html # AI # LLM # OpenSource

  13. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    🧠 新模型:Nemotron 3 Ultra(免费)(NVIDIA)上下文:1M tokens · 免费/开源 所有最新模型,每小时跟踪:https:// opensourceai.tech/latest.html # A

    🧠 New model: Nemotron 3 Ultra (free) (NVIDIA) Context: 1M tokens · free / open All the latest models, tracked hourly: https:// opensourceai.tech/latest.html # AI # LLM # OpenSource

  14. r/LocalLLaMA TIER_1 English(EN) · /u/Blahblahblakha ·

    Hy3 (295B MoE) 和 NVIDIA Nemotron-Labs-Audex-30B-A3B (支持音频的 30B MoE) GGUF 量化模型

    <!-- SC_OFF --><div class="md"><p>Sharing two GGUF quant sets, both with the same treatment: imatrix quantization, KLD/PPL measured against BF16 reference logits, llama-bench throughput numbers, and all raw benchmark data included in the repos. No vibes-based &quot;quality tested…

  15. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    🧠 新模型:Nemotron 3.5 Content Safety(免费)(NVIDIA)上下文:128k token · 免费/开源 所有最新模型,每小时跟踪:https://opensourceai.tech/la

    🧠 New model: Nemotron 3.5 Content Safety (free) (NVIDIA) Context: 128k tokens · free / open All the latest models, tracked hourly: https:// opensourceai.tech/latest.html # AI # LLM # OpenSource

  16. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    NVIDIA 发布了 Nemotron-Labs-3-Puzzle-75B-A9B,这是其 Nemotron-3-Super 混合 MoE 模型的压缩变体。该模型在...

    NVIDIA has released Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant of its Nemotron-3-Super hybrid MoE model. The model delivers 2.03x server throughput at matched user throughput while reducing active parameters from 12.8B to 9.3B. It also enables eight concurrent 1M-token …

  17. r/LocalLLaMA TIER_1 English(EN) · /u/jacek2023 ·

    nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16 · Hugging Face

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1upsdmi/nvidianvidianemotronlabs3puzzle75ba9bbf16_hugging/"> <img alt="nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16 · Hugging Face" src="https://external-preview.redd.it/9KNaWiT3A0U4xGNY8hRs0D9rkm6EHuN3da…

  18. r/LocalLLaMA TIER_1 English(EN) · /u/pmttyji ·

    nvidia/Nemotron-Labs-Audex-30B-A3B · Hugging Face

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1upnm8x/nvidianemotronlabsaudex30ba3b_hugging_face/"> <img alt="nvidia/Nemotron-Labs-Audex-30B-A3B · Hugging Face" src="https://external-preview.redd.it/I6RTEmv-UD7si5r-ldUn2OnbIIuovU5Lzh5-mTaD-04.png?width=64…

  19. Mastodon — mastodon.social TIER_1 English(EN) · strike007 ·

    本周,NVIDIA的Nemotron 3 Embed也登上了检索增强生成(RTEB)排行榜的榜首,证实了嵌入优化

    This week also saw NVIDIA’s Nemotron 3 Embed claim the top spot on the Retrieval-Augmented Generation (RTEB) leaderboard, confirming that embedding optimization is the new battleground for agentic accuracy. Meanwhile, xAI’s Grok landing on Bedrock signals a major push for enterpr…

  20. Mastodon — mastodon.social TIER_1 English(EN) · devpress ·

    Nvidia 推出免费 AI 编程工具 Nemotron Ultra 3 # ai # aicodingtools # tutorial https://www. youtube.com/watch?v=-UMmjjY37VE

    Nemotron Ultra 3 is free AI for coding from Nvidia # ai # aicodingtools # tutorial https://www. youtube.com/watch?v=-UMmjjY37VE