English(EN)NVIDIA has unveiled Nemotron 3 Ultra, a 550B parameter mixture-of-experts model for long-running AI agents. https://www. marktechpost.com/2026/06/04/nv idia-ai-
NVIDIA's new Nemotron3 Ultra is defeated by Kimi K2.6 & GLM5.1 on coding tasks like TerminalBench, etc. In order to make the Global Nemotron Coalition training committee train frontier open models, Jensen should invite at least one of the following frontier ai labs to the htt…
X — Together (inference / OSS)
TIER_1English(EN)·togethercompute·
Nemotron 3.5 ASR is built for streaming multilingual speech recognition and voice agents.
One 0.6B checkpoint. 40 language-locales. Sub-100ms latency. Cache-aware FastConformer carries context forward between chunks instead of repeatedly reprocessing overlapping audio.
Try it:
X — Together (inference / OSS)
TIER_1English(EN)·togethercompute·
Nemotron 3 Ultra is built for long-horizon agents that need to plan, code, test, debug, and iterate across large working sets.
550B parameters, 55B active, 1M context, Hybrid Mamba-Transformer MoE, Multi-Token Prediction, and multi-environment RL training.
Try it:
X — Together (inference / OSS)
TIER_1English(EN)·togethercompute·
Introducing two NVIDIA Nemotron models on Together AI: Nemotron 3 Ultra for high-throughput agentic workloads and Nemotron 3.5 ASR for low-latency multilingual speech recognition.
AI natives can now build coding agents, deep research agents, and real-time voice systems on the AI…
<p>NVIDIA released Nemotron 3.5 ASR, a cache-aware 600M streaming model transcribing 40 language-locales in real time from one checkpoint.</p> <p>The post <a href="https://www.marktechpost.com/2026/06/06/nvidia-releases-nemotron-3-5-asr-a-600m-parameter-cache-aware-streaming-mode…
<p>NVIDIA has released Nemotron 3 Ultra, a 550B total (55B active) open Mixture-of-Experts hybrid Mamba-Transformer for long-running agents. It pairs a 1M-token context with up to ~6x higher inference throughput than comparable open LLMs at on-par accuracy, and ships with open we…
NVIDIA has released Nemotron 3.5 ASR, a 600M-parameter streaming speech recognition model that transcribes 40 language-locales in real time from a single checkpoint. The cache-aware FastConformer-RNNT architecture processes each audio frame once, achieving 17x the concurrent stre…
<!-- SC_OFF --><div class="md"><p>I was originally using Parakeet for my speech recognition pipeline but decided to give Nemotron 3.5 a shot. After</p> <p>testing it on some multilingual audio clips, it's been working great so far.</p> <p>What sold me:</p> <p>- Better language su…
NVIDIA has unveiled Nemotron 3 Ultra, a 550B parameter mixture-of-experts model for long-running AI agents. https://www. marktechpost.com/2026/06/04/nv idia-ai-releases-nemotron-3-ultra-an-open-550b-mixture-of-experts-hybrid-mamba-transformer-for-long-running-agents/ # AIagent # …