PulseAugur
EN
LIVE 14:10:02

NVIDIA Nemotron 3.5 Lightning now available on Together AI for agent workloads

Together AI has announced the availability of NVIDIA's Nemotron 3.5 Lightning model. This open model is designed for high-volume, specialized agent workloads, offering up to 4x higher throughput and 30% faster task completion. It features a 30B hybrid MoE architecture with 3B active parameters, a 1M context window, and DFlash speculative decoding, making it a fast and customizable option for developers. AI

IMPACT Provides a fast, open-source model for high-volume agent tasks, potentially accelerating development in specialized AI applications.

RANK_REASON Frontier-lab model release with system card.

Read on X — Together (inference / OSS) →

AI-generated summary · Google Gemini · from 4 sources. How we write summaries →

NVIDIA Nemotron 3.5 Lightning now available on Together AI for agent workloads

COVERAGE [4]

  1. NVIDIA Blog TIER_1 English(EN) · Kari Briski ·

    NVIDIA Nemotron 3.5 Lightning and NeMo Switchyard Deliver Faster, Smarter, More Efficient Agentic AI

    As AI shifts from chatbots to autonomous agents, open models are serving market demands for full control over where AI runs and how it’s deployed and evolves. Today, NVIDIA is expanding its Nemotron 3 model family with Nemotron 3.5 Lightning, the highest-efficiency model in its c…

  2. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    Nemotron 3.5 Lightning is available on Together AI through Dedicated Model Inference, giving teams reserved capacity and predictable performance for high-volume

    Nemotron 3.5 Lightning is available on Together AI through Dedicated Model Inference, giving teams reserved capacity and predictable performance for high-volume agent workloads. Start building: https://t.co/4U09N6nL56

  3. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    What Nemotron 3.5 Lightning brings:

    What Nemotron 3.5 Lightning brings: → Fastest open model in its class → Up to 4x higher throughput → Up to 30% faster task completion → 30B hybrid MoE (3B active) → 1M context and DFlash speculative decoding → Open and customizable for post-training

  4. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    NVIDIA Nemotron 3.5 Lightning is now live on Together AI.

    NVIDIA Nemotron 3.5 Lightning is now live on Together AI. The fastest open model in its class is built for always-on agents that need to complete high-volume, specialized work quickly. https://t.co/ScHmJGOIGk