PulseAugur
EN
LIVE 18:10:59

NVIDIA releases open 550B Nemotron 3 models for agents and ASR

NVIDIA has released its Nemotron 3 family of open-source models, including Nemotron 3 Ultra and Nemotron 3.5 ASR. Nemotron 3 Ultra is a 550 billion parameter model designed for long-running AI agents, featuring a hybrid Mamba-Transformer architecture and a 1 million token context window. Nemotron 3.5 ASR is optimized for streaming speech recognition and voice agents. These models are available on Together AI, offering high inference throughput and accuracy for various AI applications. AI

IMPACT These open-source models, particularly Nemotron 3 Ultra with its large context and agent focus, could accelerate development of sophisticated AI agents and real-time voice applications.

RANK_REASON NVIDIA released new open-source models with detailed technical specifications.

Read on Mastodon — fosstodon.org →

AI-generated summary · Google Gemini · from 11 sources. How we write summaries →

NVIDIA releases open 550B Nemotron 3 models for agents and ASR

COVERAGE [11]

  1. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    NVIDIA's new Nemotron3 Ultra is defeated by Kimi K2.6 & GLM5.1 on coding tasks like TerminalBench, etc. In order to make the Global Nemotron Coalition train

    NVIDIA's new Nemotron3 Ultra is defeated by Kimi K2.6 & GLM5.1 on coding tasks like TerminalBench, etc. In order to make the Global Nemotron Coalition training committee train frontier open models, Jensen should invite at least one of the following frontier ai labs to the htt…

  2. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    Nemotron 3.5 ASR is built for streaming multilingual speech recognition and voice agents.

    Nemotron 3.5 ASR is built for streaming multilingual speech recognition and voice agents. One 0.6B checkpoint. 40 language-locales. Sub-100ms latency. Cache-aware FastConformer carries context forward between chunks instead of repeatedly reprocessing overlapping audio. Try it:

  3. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    Nemotron 3 Ultra is built for long-horizon agents that need to plan, code, test, debug, and iterate across large working sets.

    Nemotron 3 Ultra is built for long-horizon agents that need to plan, code, test, debug, and iterate across large working sets. 550B parameters, 55B active, 1M context, Hybrid Mamba-Transformer MoE, Multi-Token Prediction, and multi-environment RL training. Try it:

  4. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    Introducing two NVIDIA Nemotron models on Together AI: Nemotron 3 Ultra for high-throughput agentic workloads and Nemotron 3.5 ASR for low-latency multilingual

    Introducing two NVIDIA Nemotron models on Together AI: Nemotron 3 Ultra for high-throughput agentic workloads and Nemotron 3.5 ASR for low-latency multilingual speech recognition. AI natives can now build coding agents, deep research agents, and real-time voice systems on the AI…

  5. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    NVIDIA Releases Nemotron 3.5 ASR: A 600M-Parameter Cache-Aware Streaming Model Transcribing 40 Language-Locales in Real Time

    <p>NVIDIA released Nemotron 3.5 ASR, a cache-aware 600M streaming model transcribing 40 language-locales in real time from one checkpoint.</p> <p>The post <a href="https://www.marktechpost.com/2026/06/06/nvidia-releases-nemotron-3-5-asr-a-600m-parameter-cache-aware-streaming-mode…

  6. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    NVIDIA AI Releases Nemotron 3 Ultra: An Open 550B Mixture-of-Experts Hybrid Mamba-Transformer for Long-Running Agents

    <p>NVIDIA has released Nemotron 3 Ultra, a 550B total (55B active) open Mixture-of-Experts hybrid Mamba-Transformer for long-running agents. It pairs a 1M-token context with up to ~6x higher inference throughput than comparable open LLMs at on-par accuracy, and ships with open we…

  7. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    NVIDIA has released Nemotron 3.5 ASR, a 600M-parameter streaming speech recognition model that transcribes 40 language-locales in real time from a single checkp

    NVIDIA has released Nemotron 3.5 ASR, a 600M-parameter streaming speech recognition model that transcribes 40 language-locales in real time from a single checkpoint. The cache-aware FastConformer-RNNT architecture processes each audio frame once, achieving 17x the concurrent stre…

  8. Mastodon — fosstodon.org TIER_1 日本語(JA) · [email protected] ·

    NVIDIA, free 550B agent LLM 'Nemotron 3 Ultra' with 5x inference speed – PC Watch https://www.yayafa.com/2817421/ #AgenticAi #AI #ArtificialGeneralIntelligence #ArtificialIntel

    NVIDIA、推論5倍速で無償の550Bエージェント向けLLM「Nemotron 3 Ultra」 – PC Watch https://www. yayafa.com/2817421/ # AgenticAi # AI # ArtificialGeneralIntelligence # ArtificialIntelligence # NVIDIA # エージェント型AI # その他 # 人工知能 # 市場 # 汎用人工知能

  9. r/LocalLLaMA TIER_1 English(EN) · /u/Apart_Boat9666 ·

    Dockerized Nemotron 3.5 ASR — Switched from Parakeet, better multilingual support + streaming (4.5x realtime speed on cpu)

    <!-- SC_OFF --><div class="md"><p>I was originally using Parakeet for my speech recognition pipeline but decided to give Nemotron 3.5 a shot. After</p> <p>testing it on some multilingual audio clips, it's been working great so far.</p> <p>What sold me:</p> <p>- Better language su…

  10. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    NVIDIA has unveiled Nemotron 3 Ultra, a 550B parameter mixture-of-experts model for long-running AI agents. https://www. marktechpost.com/2026/06/04/nv idia-ai-

    NVIDIA has unveiled Nemotron 3 Ultra, a 550B parameter mixture-of-experts model for long-running AI agents. https://www. marktechpost.com/2026/06/04/nv idia-ai-releases-nemotron-3-ultra-an-open-550b-mixture-of-experts-hybrid-mamba-transformer-for-long-running-agents/ # AIagent # …

  11. Mastodon — mastodon.social TIER_1 日本語(JA) · [email protected] ·

    NVIDIA's free 550B agent LLM "Nemotron 3 Ultra" with 5x inference speed https:// pc.watch.impress.co.jp/docs/ne ws/2114675.html # impress # market # AI # Other

    NVIDIA、推論5倍速で無償の550Bエージェント向けLLM「Nemotron 3 Ultra」 https:// pc.watch.impress.co.jp/docs/ne ws/2114675.html # impress # 市場 # AI # その他