Alibaba has released and open-sourced Qwen3.8-Flash-Next, a multimodal Mixture-of-Experts model that previews the upcoming Qwen4 architecture. This new model boasts 125 billion total parameters but activates only 6 billion per token, significantly reducing training and inference costs. Qwen3.8-Flash-Next reportedly achieves performance surpassing competitors like Claude Opus 4.6 and DeepSeek-V4-Flash on various benchmarks, including coding and office tasks, while offering substantially lower API pricing. AI
IMPACT Sets a new benchmark for cost-efficiency in LLMs, potentially pressuring competitors like OpenAI and Anthropic on pricing and performance.
RANK_REASON Frontier-lab model release with system card.
- Alibaba Group
- Anthropic
- Claude Opus 4.6
- DeepSeek-V4-Flash
- OpenAI
- Qwen3.8-Flash-Next
- Qwen4
- Claude Opus 4.8
- Qwen3.7-Plus
- Gated Residual (GR)
- N-gram Embedding
- Qwen3.8-Flash
- Qwen Sparse Attention (QSA)
AI-generated summary · Google Gemini · from 9 sources. How we write summaries →