NVIDIA releases compressed Nemotron LLMs with enhanced audio and inference capabilities · 10 sources tracked
ByPulseAugur Editorial·[20 sources]·
NVIDIA has released several new models based on its Nemotron architecture, including the Nemotron-Labs-Audex-2B and Nemotron-Labs-Audex-30B-A3B, which are unified audio-text large language models. Additionally, the Nemotron-Labs-3-Puzzle-75B-A9B is a deployment-optimized, compressed variant of the Nemotron-3-Super-120B-A12B, designed for improved inference efficiency and throughput. These models are being integrated with frameworks like LangChain and Fireworks for advanced agentic capabilities, offering significant performance gains and cost reductions.
AI
IMPACT
These releases enhance audio processing capabilities and significantly improve inference efficiency, potentially lowering costs for deploying advanced LLMs.
RANK_REASON
Multiple new model releases from a frontier lab (NVIDIA), including audio-text capabilities and compressed variants for efficiency.
Our new cookbook implements Ralph loops with @nvidia Nemotron 3 Ultra via @LangChain Deep Agents: fresh context every iteration using the filesystem as memory between iterations.
Full notebook: https://t.co/j6BR187vN5 https://t.co/0ZdOyPA50I
X — Fireworks (inference infra)
TIER_1English(EN)·FireworksAI_HQ·
ICYMI: @LangChain Deep Agents on @nvidia Nemotron 3 Ultra. frontier open-model agents at ~10x lower cost than closed.
Run on Fireworks, then post-train it into specialized intelligence you own.
https://t.co/IiMMr94QLa https://t.co/RZoQp40iTk
<p>NVIDIA released Nemotron 3 Embed on July 15 and 16, 2026. The collection has three open checkpoints: Nemotron-3-Embed-8B-BF16, Nemotron-3-Embed-1B-BF16, and Nemotron-3-Embed-1B-NVFP4. The 8B ranks #1 on RTEB at 78.46 average NDCG@10. The 1B came from ModelOpt NAS pruning plus …
<p>NVIDIA has released Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant of Nemotron-3-Super. Iterative Puzzle alternates hardware-aware structural compression with short knowledge distillation recovery phases. The model drops from 120.7B total / 12.8B active parameters to 75.…
NVIDIA has unveiled Nemotron 3 Embed, an open-source embedding model collection. The 8 billion parameter version tops the RTEB benchmark for retrieval tasks, marking a significant advancement in AI infrastructure for enterprise search and RAG applications. https://www. marktechpo…
<h2> What Changed </h2> <p>NVIDIA has introduced Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4, a new large language model (LLM) designed for enhanced inference efficiency. This model is a compressed variant of the Nemotron-3-Super-120B-A12B, specifically optimized for deployment in inter…
🧠 New model: Nemotron 3 Ultra (NVIDIA) Context: 1M tokens · $0.5 in / $2.2 out per M All the latest models, tracked hourly: https:// opensourceai.tech/latest.html # AI # LLM # OpenSource
<!-- SC_OFF --><div class="md"><p>Sharing two GGUF quant sets, both with the same treatment: imatrix quantization, KLD/PPL measured against BF16 reference logits, llama-bench throughput numbers, and all raw benchmark data included in the repos. No vibes-based "quality tested…
NVIDIA has released Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant of its Nemotron-3-Super hybrid MoE model. The model delivers 2.03x server throughput at matched user throughput while reducing active parameters from 12.8B to 9.3B. It also enables eight concurrent 1M-token …
This week also saw NVIDIA’s Nemotron 3 Embed claim the top spot on the Retrieval-Augmented Generation (RTEB) leaderboard, confirming that embedding optimization is the new battleground for agentic accuracy. Meanwhile, xAI’s Grok landing on Bedrock signals a major push for enterpr…