PulseAugur
EN
LIVE 05:00:50

NVIDIA releases compressed Nemotron LLMs with enhanced audio and inference capabilities · 10 sources tracked

NVIDIA has released several new models based on its Nemotron architecture, including the Nemotron-Labs-Audex-2B and Nemotron-Labs-Audex-30B-A3B, which are unified audio-text large language models. Additionally, the Nemotron-Labs-3-Puzzle-75B-A9B is a deployment-optimized, compressed variant of the Nemotron-3-Super-120B-A12B, designed for improved inference efficiency and throughput. These models are being integrated with frameworks like LangChain and Fireworks for advanced agentic capabilities, offering significant performance gains and cost reductions. AI

IMPACT These releases enhance audio processing capabilities and significantly improve inference efficiency, potentially lowering costs for deploying advanced LLMs.

RANK_REASON Multiple new model releases from a frontier lab (NVIDIA), including audio-text capabilities and compressed variants for efficiency.

Read on Hugging Face Trending Models →

AI-generated summary · Google Gemini · from 20 sources. How we write summaries →

NVIDIA releases compressed Nemotron LLMs with enhanced audio and inference capabilities · 10 sources tracked

COVERAGE [20]

  1. Hugging Face Blog TIER_1 English(EN) ·

    NVIDIA Nemotron 3 Embed Ranks #1 Overall on RTEB, Advancing Agentic Retrieval

  2. Hugging Face Trending Models TIER_1 (CA) · nvidia ·

    nvidia/Nemotron-Labs-Audex-2B

    text-generation · 1,050 downloads · 50 likes

  3. Hugging Face Trending Models TIER_1 English(EN) · nvidia ·

    nvidia/Nemotron-Labs-Audex-30B-A3B

    text-generation · 0 downloads · 46 likes

  4. Hugging Face Trending Models TIER_1 Italiano(IT) · nvidia ·

    nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4

    text-generation · 47 downloads · 53 likes

  5. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    Learn more about the LangChain Deep Agents Harness optimized for NVIDIA Nemotron 3 Ultra: https://t.co/w1aOqdBazc

    Learn more about the LangChain Deep Agents Harness optimized for NVIDIA Nemotron 3 Ultra: https://t.co/w1aOqdBazc

  6. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    Our new cookbook implements Ralph loops with @nvidia Nemotron 3 Ultra via @LangChain Deep Agents: fresh context every iteration using the filesystem as memory b

    Our new cookbook implements Ralph loops with @nvidia Nemotron 3 Ultra via @LangChain Deep Agents: fresh context every iteration using the filesystem as memory between iterations. Full notebook: https://t.co/j6BR187vN5 https://t.co/0ZdOyPA50I

  7. X — Fireworks (inference infra) TIER_1 English(EN) · FireworksAI_HQ ·

    ICYMI: @LangChain Deep Agents on @nvidia Nemotron 3 Ultra. frontier open-model agents at ~10x lower cost than closed.

    ICYMI: @LangChain Deep Agents on @nvidia Nemotron 3 Ultra. frontier open-model agents at ~10x lower cost than closed. Run on Fireworks, then post-train it into specialized intelligence you own. https://t.co/IiMMr94QLa https://t.co/RZoQp40iTk

  8. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    NVIDIA AI Releases Nemotron 3 Embed: An Open Embedding Collection Whose 8B Checkpoint Ranks #1 on RTEB

    <p>NVIDIA released Nemotron 3 Embed on July 15 and 16, 2026. The collection has three open checkpoints: Nemotron-3-Embed-8B-BF16, Nemotron-3-Embed-1B-BF16, and Nemotron-3-Embed-1B-NVFP4. The 8B ranks #1 on RTEB at 78.46 average NDCG@10. The 1B came from ModelOpt NAS pruning plus …

  9. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    NVIDIA Releases Nemotron-Labs-3-Puzzle-75B-A9B: A Compressed Hybrid MoE LLM Delivering 2.03x Server Throughput at Matched User Throughput

    <p>NVIDIA has released Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant of Nemotron-3-Super. Iterative Puzzle alternates hardware-aware structural compression with short knowledge distillation recovery phases. The model drops from 120.7B total / 12.8B active parameters to 75.…

  10. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    NVIDIA has unveiled Nemotron 3 Embed, an open-source embedding model collection. The 8 billion parameter version tops the RTEB benchmark for retrieval tasks, ma

    NVIDIA has unveiled Nemotron 3 Embed, an open-source embedding model collection. The 8 billion parameter version tops the RTEB benchmark for retrieval tasks, marking a significant advancement in AI infrastructure for enterprise search and RAG applications. https://www. marktechpo…

  11. dev.to — LLM tag TIER_1 English(EN) · Pneumetron ·

    NVIDIA Unveils Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4: A Deployment-Optimized Hybrid MoE LLM

    <h2> What Changed </h2> <p>NVIDIA has introduced Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4, a new large language model (LLM) designed for enhanced inference efficiency. This model is a compressed variant of the Nemotron-3-Super-120B-A12B, specifically optimized for deployment in inter…

  12. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    🧠 New model: Nemotron 3 Ultra (NVIDIA) Context: 1M tokens · $0.5 in / $2.2 out per M All the latest models, tracked hourly: https:// opensourceai.tech/latest.ht

    🧠 New model: Nemotron 3 Ultra (NVIDIA) Context: 1M tokens · $0.5 in / $2.2 out per M All the latest models, tracked hourly: https:// opensourceai.tech/latest.html # AI # LLM # OpenSource

  13. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    🧠 New model: Nemotron 3 Ultra (free) (NVIDIA) Context: 1M tokens · free / open All the latest models, tracked hourly: https:// opensourceai.tech/latest.html # A

    🧠 New model: Nemotron 3 Ultra (free) (NVIDIA) Context: 1M tokens · free / open All the latest models, tracked hourly: https:// opensourceai.tech/latest.html # AI # LLM # OpenSource

  14. r/LocalLLaMA TIER_1 English(EN) · /u/Blahblahblakha ·

    Hy3 (295B MoE) and NVIDIA Nemotron-Labs-Audex-30B-A3B (audio-capable 30B MoE) GGUF quants

    <!-- SC_OFF --><div class="md"><p>Sharing two GGUF quant sets, both with the same treatment: imatrix quantization, KLD/PPL measured against BF16 reference logits, llama-bench throughput numbers, and all raw benchmark data included in the repos. No vibes-based &quot;quality tested…

  15. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    🧠 New model: Nemotron 3.5 Content Safety (free) (NVIDIA) Context: 128k tokens · free / open All the latest models, tracked hourly: https:// opensourceai.tech/la

    🧠 New model: Nemotron 3.5 Content Safety (free) (NVIDIA) Context: 128k tokens · free / open All the latest models, tracked hourly: https:// opensourceai.tech/latest.html # AI # LLM # OpenSource

  16. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    NVIDIA has released Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant of its Nemotron-3-Super hybrid MoE model. The model delivers 2.03x server throughput at

    NVIDIA has released Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant of its Nemotron-3-Super hybrid MoE model. The model delivers 2.03x server throughput at matched user throughput while reducing active parameters from 12.8B to 9.3B. It also enables eight concurrent 1M-token …

  17. r/LocalLLaMA TIER_1 English(EN) · /u/jacek2023 ·

    nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16 · Hugging Face

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1upsdmi/nvidianvidianemotronlabs3puzzle75ba9bbf16_hugging/"> <img alt="nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16 · Hugging Face" src="https://external-preview.redd.it/9KNaWiT3A0U4xGNY8hRs0D9rkm6EHuN3da…

  18. r/LocalLLaMA TIER_1 English(EN) · /u/pmttyji ·

    nvidia/Nemotron-Labs-Audex-30B-A3B · Hugging Face

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1upnm8x/nvidianemotronlabsaudex30ba3b_hugging_face/"> <img alt="nvidia/Nemotron-Labs-Audex-30B-A3B · Hugging Face" src="https://external-preview.redd.it/I6RTEmv-UD7si5r-ldUn2OnbIIuovU5Lzh5-mTaD-04.png?width=64…

  19. Mastodon — mastodon.social TIER_1 English(EN) · strike007 ·

    This week also saw NVIDIA’s Nemotron 3 Embed claim the top spot on the Retrieval-Augmented Generation (RTEB) leaderboard, confirming that embedding optimization

    This week also saw NVIDIA’s Nemotron 3 Embed claim the top spot on the Retrieval-Augmented Generation (RTEB) leaderboard, confirming that embedding optimization is the new battleground for agentic accuracy. Meanwhile, xAI’s Grok landing on Bedrock signals a major push for enterpr…

  20. Mastodon — mastodon.social TIER_1 English(EN) · devpress ·

    Nemotron Ultra 3 is free AI for coding from Nvidia # ai # aicodingtools # tutorial https://www. youtube.com/watch?v=-UMmjjY37VE

    Nemotron Ultra 3 is free AI for coding from Nvidia # ai # aicodingtools # tutorial https://www. youtube.com/watch?v=-UMmjjY37VE