NVIDIA发布了基于其Nemotron架构的几款新模型,包括统一的音频-文本大语言模型Nemotron-Labs-Audex-2B和Nemotron-Labs-Audex-30B-A3B。此外,Nemotron-Labs-3-Puzzle-75B-A9B是Nemotron-3-Super-120B-A12B的部署优化压缩版本,旨在提高推理效率和吞吐量。这些模型正与LangChain和Fireworks等框架集成,以实现高级代理能力,提供显著的性能提升和成本降低。
AI
Our new cookbook implements Ralph loops with @nvidia Nemotron 3 Ultra via @LangChain Deep Agents: fresh context every iteration using the filesystem as memory between iterations.
Full notebook: https://t.co/j6BR187vN5 https://t.co/0ZdOyPA50I
X — Fireworks (inference infra)
TIER_1English(EN)·FireworksAI_HQ·
ICYMI: @LangChain Deep Agents on @nvidia Nemotron 3 Ultra. frontier open-model agents at ~10x lower cost than closed.
Run on Fireworks, then post-train it into specialized intelligence you own.
https://t.co/IiMMr94QLa https://t.co/RZoQp40iTk
<p>NVIDIA released Nemotron 3 Embed on July 15 and 16, 2026. The collection has three open checkpoints: Nemotron-3-Embed-8B-BF16, Nemotron-3-Embed-1B-BF16, and Nemotron-3-Embed-1B-NVFP4. The 8B ranks #1 on RTEB at 78.46 average NDCG@10. The 1B came from ModelOpt NAS pruning plus …
<p>NVIDIA has released Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant of Nemotron-3-Super. Iterative Puzzle alternates hardware-aware structural compression with short knowledge distillation recovery phases. The model drops from 120.7B total / 12.8B active parameters to 75.…
NVIDIA has unveiled Nemotron 3 Embed, an open-source embedding model collection. The 8 billion parameter version tops the RTEB benchmark for retrieval tasks, marking a significant advancement in AI infrastructure for enterprise search and RAG applications. https://www. marktechpo…
<h2> What Changed </h2> <p>NVIDIA has introduced Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4, a new large language model (LLM) designed for enhanced inference efficiency. This model is a compressed variant of the Nemotron-3-Super-120B-A12B, specifically optimized for deployment in inter…
🧠 New model: Nemotron 3 Ultra (NVIDIA) Context: 1M tokens · $0.5 in / $2.2 out per M All the latest models, tracked hourly: https:// opensourceai.tech/latest.html # AI # LLM # OpenSource
<!-- SC_OFF --><div class="md"><p>Sharing two GGUF quant sets, both with the same treatment: imatrix quantization, KLD/PPL measured against BF16 reference logits, llama-bench throughput numbers, and all raw benchmark data included in the repos. No vibes-based "quality tested…
NVIDIA has released Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant of its Nemotron-3-Super hybrid MoE model. The model delivers 2.03x server throughput at matched user throughput while reducing active parameters from 12.8B to 9.3B. It also enables eight concurrent 1M-token …
This week also saw NVIDIA’s Nemotron 3 Embed claim the top spot on the Retrieval-Augmented Generation (RTEB) leaderboard, confirming that embedding optimization is the new battleground for agentic accuracy. Meanwhile, xAI’s Grok landing on Bedrock signals a major push for enterpr…