NVIDIA has released new draft models under the Nemotron 3.5 Lightning 30B-A3B series, designed for specialized decoding tasks. Nemotron-3.5-Lightning-30B-A3B-NVFP4-DFlash, with 833 million parameters, accelerates a 30B target model for DFlash decoding on hardware like H100 and RTX 5090. Another variant, Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark, features 967 million parameters and optimizes latency for DSpark decoding on DGX Spark and H100 systems. These models utilize a hybrid Mixture-of-Experts architecture combining Mamba2 and Transformer components, with 3 billion active parameters, and support a 1 million token context window under the OpenMDW-1.1 license. AI
IMPACT These specialized draft models may offer performance improvements for specific decoding tasks and hardware configurations.
RANK_REASON New model release from a frontier lab (NVIDIA).
Read on Mastodon — mastodon.social →
- DGX Spark
- DSpark
- FLASH
- Mamba2
- Nemotron 3.5 Lightning 30B-A3B
- Nemotron-3.5-Lightning-30B-A3B-NVFP4-DFlash
- Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark
- NVFP4
- NVIDIA
- Transformer
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →