The Qwen team has released the weights for their Qwen3.8-Flash-Next model, a multimodal Mixture-of-Experts (MoE) architecture. This new model incorporates innovations such as Gated DeltaNet+Qwen Sparse Attention (GDN+QSA) hybrid, Gated Residual Networks, N-gram embeddings, and the Muon optimizer to enhance performance while reducing computational and training costs. The model supports a base context of 262,144 tokens, expandable to 1 million tokens via YaRN, and its weights are available on Hugging Face and ModelScope. AI
IMPACT This release offers enhanced efficiency and a significantly larger context window, potentially enabling more complex multimodal applications.
RANK_REASON Frontier-lab model release with system card. [lever_c_demoted from frontier_release: ic=1 ai=1.0]
Read on Mastodon — fosstodon.org →
- Gated DeltaNet+Qwen Sparse Attention
- Gated Residual Networks with Dilated Convolutions for Monaural Speech Enhancement
- GDN+QSA
- Hugging Face
- ModelScope
- muon
- N-gram Embedding
- Qwen
- Qwen3.8-Flash-Next
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →