PulseAugur
EN
LIVE 13:40:26
Deutsch(DE) NVIDIA Nemotron-3.5-Lightning-30B-A3B-NVFP4-DFlash ist ein 833M-Parameter-Draft-Modell für DFlash-spezifische Dekodierung. Es beschleunigt das 30B-Zielmodell au

NVIDIA releases Nemotron 3.5 Lightning draft models for specialized decoding · 3 sources tracked

NVIDIA has released new draft models under the Nemotron 3.5 Lightning 30B-A3B series, designed for specialized decoding tasks. Nemotron-3.5-Lightning-30B-A3B-NVFP4-DFlash, with 833 million parameters, accelerates a 30B target model for DFlash decoding on hardware like H100 and RTX 5090. Another variant, Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark, features 967 million parameters and optimizes latency for DSpark decoding on DGX Spark and H100 systems. These models utilize a hybrid Mixture-of-Experts architecture combining Mamba2 and Transformer components, with 3 billion active parameters, and support a 1 million token context window under the OpenMDW-1.1 license. AI

IMPACT These specialized draft models may offer performance improvements for specific decoding tasks and hardware configurations.

RANK_REASON New model release from a frontier lab (NVIDIA).

Read on Mastodon — mastodon.social →

AI-generated summary · Google Gemini · from 3 sources. How we write summaries →

NVIDIA releases Nemotron 3.5 Lightning draft models for specialized decoding · 3 sources tracked

COVERAGE [3]

  1. Mastodon — mastodon.social TIER_1 Deutsch(DE) · aisyndicate ·

    NVIDIA Nemotron-3.5-Lightning-30B-A3B-NVFP4-DFlash is an 833M parameter draft model for DFlash-specific decoding. It accelerates the 30B target model

    NVIDIA Nemotron-3.5-Lightning-30B-A3B-NVFP4-DFlash ist ein 833M-Parameter-Draft-Modell für DFlash-spezifische Dekodierung. Es beschleunigt das 30B-Zielmodell auf H100/RTX 5090. SPEED-Bench-Akzeptanzrate: 3.16 (Draft-Länge 7). Lizenz: OpenMDW-1.1. https:// huggingface.co/nvidia/NV…

  2. Mastodon — mastodon.social TIER_1 Deutsch(DE) · aisyndicate ·

    NVIDIA Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark is a 967M parameter draft model for DSpark speculative decoding. It optimizes latency on DGX Spark

    NVIDIA Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark ist ein 967M-Parameter-Draft-Modell für DSpark-spekulative Dekodierung. Es optimiert die Latenz auf DGX Spark (GB10) und H100. SPEED-Bench-Akzeptanzrate: 3,75 (Draft-Länge 7). Lizenz: OpenMDW-1.1. https:// huggingface.co/nvidia/N…

  3. Mastodon — mastodon.social TIER_1 English(EN) · aisyndicate ·

    NVIDIA Nemotron 3.5 Lightning 30B-A3B: Hybrid MoE (Mamba2/Transformer) mit 3B aktiven Parametern. NVFP4-Training und Multi-Token Prediction (MTP) ermöglichen na

    NVIDIA Nemotron 3.5 Lightning 30B-A3B: Hybrid MoE (Mamba2/Transformer) mit 3B aktiven Parametern. NVFP4-Training und Multi-Token Prediction (MTP) ermöglichen native spekulative Dekodierung. 1M Kontextfenster, OpenMDW-Lizenz. https:// huggingface.co/nvidia/NVIDIA-N emotron-3.5-Lig…