Alibaba previews Qwen4 architecture with cost-efficient Qwen3.8-Flash-Next model
ByPulseAugur Editorial·[13 sources]·
Alibaba's Qwen team has released Qwen3.8-Flash-Next, an open-weight multimodal MoE model that previews the architecture for the upcoming Qwen4. This new model boasts significant cost-efficiency, activating only 6B parameters per token from a 125B backbone, with an additional 51B N-gram embeddings. It demonstrates strong performance across various benchmarks, including coding and multimodal tasks, while offering a native 262K context window extendable to 1M tokens. The architecture introduces innovations like Qwen Sparse Attention (QSA) and Gated Residuals to improve efficiency and stability.
AI
IMPACT
Sets a new standard for cost-efficiency in frontier-class models, potentially accelerating adoption in resource-constrained environments.
RANK_REASON
Frontier-lab model release with system card and open weights.
With only 6B active parameters, Qwen3.8-Flash-Next-Base tops 8 of 14 benchmarks, including MMLU-Pro, SuperGPQA, BBH and GSM8K. And it remains competitive with Qwen3.7-Plus-Base on the rest. Its 51B N-gram embedding parameters use deterministic lookups, adding no per-token https:/…
X — Qwen (Alibaba)
TIER_1English(EN)·Alibaba_Qwen·
At a 1M-token context length, QSA’s attention kernel is up to 7.6× faster in prefill and 4.9× faster in decode. With a 90% prefix-cache hit rate, Qwen3.8-Flash-Next delivers 8.6× the prefill throughput of Qwen3.7-Plus. https://t.co/QZ0koCWvQU
X — Qwen (Alibaba)
TIER_1English(EN)·Alibaba_Qwen·
Model Architecture
Four core upgrades for maximum capability, efficiency, capacity, and stability:
- Attention: GDN + QSA Hybrid. Gated DeltaNet (GDN) compresses history. Qwen Sparse Attention (QSA) uses a lightweight indexer for micro-block context selection. Lower the cost of…
X — Qwen (Alibaba)
TIER_1English(EN)·Alibaba_Qwen·
⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight!
The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens.
125B parameters + 51B N-gram https://…
In this release we are opening the weights of Qwen3.8-Flash-Next, a multimodal MoE model that also serves as an early preview of the architecture used in Qwen4. It plays the same role that Qwen3-Next played for Qwen3.5: the hybrid Gated DeltaNet + Gated Attention design introduce…
Hacker News — AI stories ≥50 points
TIER_1English(EN)·tosh·
<p>We look at Qwen3.8-Flash-Next, Alibaba's open-weight multimodal Mixture-of-Experts model and an early preview of the Qwen4 architecture. We break down where the 180B parameters actually sit: a 125B backbone, a 51B N-gram embedding table, and a 4B multi-token prediction module,…
dev.to — LLM tag
TIER_1English(EN)·James Anderson·
<p>Here's a number that shouldn't make sense on first read.</p> <p>Qwen just released Qwen3.8-Flash-Next, and it activates <strong>6 billion parameters per token</strong> — while matching or beating models that activate <strong>13B (DeepSeek-V4-Flash) and 17B (Qwen3.7-Plus)</stro…
dev.to — LLM tag
TIER_1English(EN)·Mariano Gobea Alcoba·
<h2> Architectural Evolution: A Technical Deconstruction of Qwen3.8-Flash-Next </h2> <p>The release of Qwen3.8-Flash-Next marks a significant shift in the deployment strategies for large language models (LLMs) in high-throughput, low-latency environments. As infrastructure archit…
<h1> Qwen3.8-Flash-Next (2026): The Complete Guide to Qwen's Qwen4-Preview Architecture Model </h1> <h2> 🎯 Core Takeaways (TL;DR) </h2> <ul> <li> <strong>Qwen3.8-Flash-Next</strong> is Alibaba Qwen's open-weight preview of the architecture that will underpin <strong>Qwen4</strong…
Model Qwen3.8-Flash-Next wyznacza nową granicę opłacalności systemów AI, oferując wydajność klasy flagowej przy cenie zaledwie 0,16 USD za milion tokenów. # si # ai # sztucznainteligencja # wiadomości # informacje # technologia https:// aisight.pl/technologia/generat ywna-ai/llm/…