Nvidia has released Qwen3.8-Flash-Next, a 125-billion parameter model, as an NVFP4 checkpoint. This model incorporates Hybrid Attention and Mixture-of-Experts (MoE) architectures. It is optimized to run via vLLM on Nvidia's Blackwell B200/B300 hardware, with quantization reducing its disk size by 63% compared to BF16. AI
IMPACT Optimized model release for new hardware accelerates inference performance for large language models.
RANK_REASON Frontier-lab model release with system card. [lever_c_demoted from frontier_release: ic=1 ai=1.0]
Read on Mastodon — mastodon.social →
- Hybrid Attention based Multimodal Network for Spoken Language Classification
- NVFP4
- Nvidia
- Nvidia Blackwell B200
- Nvidia Blackwell B300
- Qwen3.8-Flash-Next
- vLLM
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →