PulseAugur
EN
LIVE 22:41:45

Alibaba previews Qwen4 architecture with cost-efficient Qwen3.8-Flash-Next model

Alibaba's Qwen team has released Qwen3.8-Flash-Next, an open-weight multimodal MoE model that previews the architecture for the upcoming Qwen4. This new model boasts significant cost-efficiency, activating only 6B parameters per token from a 125B backbone, with an additional 51B N-gram embeddings. It demonstrates strong performance across various benchmarks, including coding and multimodal tasks, while offering a native 262K context window extendable to 1M tokens. The architecture introduces innovations like Qwen Sparse Attention (QSA) and Gated Residuals to improve efficiency and stability. AI

IMPACT Sets a new standard for cost-efficiency in frontier-class models, potentially accelerating adoption in resource-constrained environments.

RANK_REASON Frontier-lab model release with system card and open weights.

Read on Qwen tech blog →

AI-generated summary · Google Gemini · from 13 sources. How we write summaries →

Alibaba previews Qwen4 architecture with cost-efficient Qwen3.8-Flash-Next model

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Frontier Release
Frontier-lab model release with system card and open weights.
Source corroboration
13 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
45 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+3 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

Full methodology in our editorial standards.

COVERAGE [13]

  1. X — Qwen (Alibaba) TIER_1 English(EN) · Alibaba_Qwen ·

    125B/6B, 1M context, multimodal, now in OpenCode Go. ⚡ Happy coding with Qwen3.8-Flash!

    125B/6B, 1M context, multimodal, now in OpenCode Go. ⚡ Happy coding with Qwen3.8-Flash!

  2. X — Qwen (Alibaba) TIER_1 English(EN) · Alibaba_Qwen ·

    With only 6B active parameters, Qwen3.8-Flash-Next-Base tops 8 of 14 benchmarks, including MMLU-Pro, SuperGPQA, BBH and GSM8K. And it remains competitive with Q

    With only 6B active parameters, Qwen3.8-Flash-Next-Base tops 8 of 14 benchmarks, including MMLU-Pro, SuperGPQA, BBH and GSM8K. And it remains competitive with Qwen3.7-Plus-Base on the rest. Its 51B N-gram embedding parameters use deterministic lookups, adding no per-token https:/…

  3. X — Qwen (Alibaba) TIER_1 English(EN) · Alibaba_Qwen ·

    At a 1M-token context length, QSA’s attention kernel is up to 7.6× faster in prefill and 4.9× faster in decode. With a 90% prefix-cache hit rate, Qwen3.8-Flash-

    At a 1M-token context length, QSA’s attention kernel is up to 7.6× faster in prefill and 4.9× faster in decode. With a 90% prefix-cache hit rate, Qwen3.8-Flash-Next delivers 8.6× the prefill throughput of Qwen3.7-Plus. https://t.co/QZ0koCWvQU

  4. X — Qwen (Alibaba) TIER_1 English(EN) · Alibaba_Qwen ·

    Model Architecture

    Model Architecture Four core upgrades for maximum capability, efficiency, capacity, and stability: - Attention: GDN + QSA Hybrid. Gated DeltaNet (GDN) compresses history. Qwen Sparse Attention (QSA) uses a lightweight indexer for micro-block context selection. Lower the cost of…

  5. X — Qwen (Alibaba) TIER_1 English(EN) · Alibaba_Qwen ·

    Language performance &amp; Vision Language performance: https://t.co/PkhyKwFmP3

    Language performance &amp; Vision Language performance: https://t.co/PkhyKwFmP3

  6. X — Qwen (Alibaba) TIER_1 English(EN) · Alibaba_Qwen ·

    ⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight!

    ⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight! The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens. 125B parameters + 51B N-gram https://…

  7. Qwen tech blog TIER_1 English(EN) · QwenTeam ·

    Qwen3.8-Flash-Next: A New Architecture, Towards Ultimate Cost-Efficiency

    In this release we are opening the weights of Qwen3.8-Flash-Next, a multimodal MoE model that also serves as an early preview of the architecture used in Qwen4. It plays the same role that Qwen3-Next played for Qwen3.5: the hybrid Gated DeltaNet + Gated Attention design introduce…

  8. Hacker News — AI stories ≥50 points TIER_1 English(EN) · tosh ·

    Qwen3.8-Flash-Next: A New Architecture, Towards Ultimate Cost-Efficiency

  9. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    Alibaba’s Qwen Team Releases Qwen3.8-Flash-Next: A 125B Multimodal MoE With 6B Active Parameters Previewing the Qwen4 Architecture

    <p>We look at Qwen3.8-Flash-Next, Alibaba's open-weight multimodal Mixture-of-Experts model and an early preview of the Qwen4 architecture. We break down where the 180B parameters actually sit: a 125B backbone, a 51B N-gram embedding table, and a 4B multi-token prediction module,…

  10. dev.to — LLM tag TIER_1 English(EN) · James Anderson ·

    How a 6B-Active Model Beats 17B-Active Ones: What Qwen3.8-Flash-Next Actually Changed

    <p>Here's a number that shouldn't make sense on first read.</p> <p>Qwen just released Qwen3.8-Flash-Next, and it activates <strong>6 billion parameters per token</strong> — while matching or beating models that activate <strong>13B (DeepSeek-V4-Flash) and 17B (Qwen3.7-Plus)</stro…

  11. dev.to — LLM tag TIER_1 English(EN) · Mariano Gobea Alcoba ·

    Qwen3.8-Flash-Next Intelligence, Performance and Price Analysis!

    <h2> Architectural Evolution: A Technical Deconstruction of Qwen3.8-Flash-Next </h2> <p>The release of Qwen3.8-Flash-Next marks a significant shift in the deployment strategies for large language models (LLMs) in high-throughput, low-latency environments. As infrastructure archit…

  12. dev.to — LLM tag TIER_1 English(EN) · cz ·

    Qwen3.8-Flash-Next (2026): The Complete Guide to Qwen’s Qwen4-Preview Architecture Model

    <h1> Qwen3.8-Flash-Next (2026): The Complete Guide to Qwen's Qwen4-Preview Architecture Model </h1> <h2> 🎯 Core Takeaways (TL;DR) </h2> <ul> <li> <strong>Qwen3.8-Flash-Next</strong> is Alibaba Qwen's open-weight preview of the architecture that will underpin <strong>Qwen4</strong…

  13. Mastodon — mastodon.social TIER_1 Polski(PL) · aisight ·

    Qwen3.8-Flash-Next Model Sets New AI System Cost-Effectiveness Benchmark, Offering Flagship Performance at Just $0.16 Per Million Tokens. # si

    Model Qwen3.8-Flash-Next wyznacza nową granicę opłacalności systemów AI, oferując wydajność klasy flagowej przy cenie zaledwie 0,16 USD za milion tokenów. # si # ai # sztucznainteligencja # wiadomości # informacje # technologia https:// aisight.pl/technologia/generat ywna-ai/llm/…