Alibaba's Qwen team has released Qwen3.8-Flash, an open-weight multimodal model that serves as an early preview of the Qwen4 architecture. This new model boasts significant improvements in cost-efficiency and performance, outperforming its predecessor Qwen3.7-Plus across various benchmarks, particularly in coding and office tasks. Key architectural upgrades include a hybrid attention mechanism, gated residual networks, and an N-gram embedding system, contributing to its enhanced capabilities and reduced computational costs. AI
IMPACT Sets new SOTA on several benchmarks with a focus on cost-efficiency, potentially influencing future model development and deployment strategies.
RANK_REASON Frontier-lab model release with system card and benchmark results.
- Alibaba Group
- GSM8K
- MMLU-Pro
- Muon optimizer
- N-gram Embedding
- Qwen
- Qwen3.7-Plus
- Qwen3.8-Flash
- Qwen3.8-Flash-Next
- Qwen4
- QwenCloud
- Qwen Sparse Attention
- SuperGPQA
AI-generated summary · Google Gemini · from 5 sources. How we write summaries →