PulseAugur
中
实时 20:46:08
English(EN) Qwen3.8-Flash-Next: A New Architecture, Towards Ultimate Cost-Efficiency

阿里巴巴预览Qwen4架构,推出成本效益高的Qwen3.8-Flash-Next模型

阿里巴巴的Qwen团队发布了Qwen3.8-Flash-Next,这是一个开放权重的多模态MoE模型,预览了即将推出的Qwen4的架构。该新模型拥有显著的成本效益,在125B的骨干网络中每个token仅激活6B参数,并额外包含51B的N-gram嵌入。它在包括编码和多模态任务在内的各种基准测试中表现强劲,同时提供原生262K上下文窗口,可扩展至1M token。该架构引入了Qwen Sparse Attention (QSA) 和 Gated Residuals 等创新,以提高效率和稳定性。 AI

影响 为前沿模型设定了新的成本效益标准,有可能加速在资源受限环境中的应用。

排序理由 前沿实验室模型发布,附带系统卡和开放权重。

在 Qwen tech blog 阅读 →

AI 生成摘要 · Google Gemini · 来自 13 个来源。 我们如何撰写摘要 →

阿里巴巴预览Qwen4架构,推出成本效益高的Qwen3.8-Flash-Next模型

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Frontier Release
前沿实验室模型发布,附带系统卡和开放权重。
Source corroboration
13 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
45 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+3 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [13]

  1. X — Qwen (Alibaba) TIER_1 English(EN) · Alibaba_Qwen ·

    125B/6B,1M上下文,多模态,现已集成到OpenCode Go。⚡ 祝您使用Qwen3.8-Flash编码愉快!

    125B/6B, 1M context, multimodal, now in OpenCode Go. ⚡ Happy coding with Qwen3.8-Flash!

  2. X — Qwen (Alibaba) TIER_1 English(EN) · Alibaba_Qwen ·

    仅用60亿活跃参数,Qwen3.8-Flash-Next-Base在MMLU-Pro、SuperGPQA、BBH和GSM8K等14项基准测试中名列前茅,并保持与Q的竞争力

    With only 6B active parameters, Qwen3.8-Flash-Next-Base tops 8 of 14 benchmarks, including MMLU-Pro, SuperGPQA, BBH and GSM8K. And it remains competitive with Qwen3.7-Plus-Base on the rest. Its 51B N-gram embedding parameters use deterministic lookups, adding no per-token https:/…

  3. X — Qwen (Alibaba) TIER_1 English(EN) · Alibaba_Qwen ·

    在100万token的上下文长度下,QSA的注意力核在预填充时速度提升高达7.6倍,在解码时速度提升高达4.9倍。在90%的前缀缓存命中率下,Qwen3.8-Flash-

    At a 1M-token context length, QSA’s attention kernel is up to 7.6× faster in prefill and 4.9× faster in decode. With a 90% prefix-cache hit rate, Qwen3.8-Flash-Next delivers 8.6× the prefill throughput of Qwen3.7-Plus. https://t.co/QZ0koCWvQU

  4. X — Qwen (Alibaba) TIER_1 English(EN) · Alibaba_Qwen ·

    模型架构

    Model Architecture Four core upgrades for maximum capability, efficiency, capacity, and stability: - Attention: GDN + QSA Hybrid. Gated DeltaNet (GDN) compresses history. Qwen Sparse Attention (QSA) uses a lightweight indexer for micro-block context selection. Lower the cost of…

  5. X — Qwen (Alibaba) TIER_1 English(EN) · Alibaba_Qwen ·

    语言性能与视觉语言性能:https://t.co/PkhyKwFmP3

    Language performance &amp; Vision Language performance: https://t.co/PkhyKwFmP3

  6. X — Qwen (Alibaba) TIER_1 English(EN) · Alibaba_Qwen ·

    ⚡认识 Qwen3.8-Flash,一个多模态 MoE 模型,也是 Qwen4 架构的早期预览版,现已开源!

    ⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight! The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens. 125B parameters + 51B N-gram https://…

  7. Qwen tech blog TIER_1 English(EN) · QwenTeam ·

    Qwen3.8-Flash-Next:新架构,迈向极致性价比

    In this release we are opening the weights of Qwen3.8-Flash-Next, a multimodal MoE model that also serves as an early preview of the architecture used in Qwen4. It plays the same role that Qwen3-Next played for Qwen3.5: the hybrid Gated DeltaNet + Gated Attention design introduce…

  8. Hacker News — AI stories ≥50 points TIER_1 English(EN) · tosh ·

    Qwen3.8-Flash-Next:新架构,迈向极致性价比

  9. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    阿里巴巴Qwen团队发布Qwen3.8-Flash-Next:一个拥有60亿活跃参数的1250亿多模态MoE模型,预览Qwen4架构

    <p>We look at Qwen3.8-Flash-Next, Alibaba's open-weight multimodal Mixture-of-Experts model and an early preview of the Qwen4 architecture. We break down where the 180B parameters actually sit: a 125B backbone, a 51B N-gram embedding table, and a 4B multi-token prediction module,…

  10. dev.to — LLM tag TIER_1 English(EN) · James Anderson ·

    一个6B活跃模型如何击败17B活跃模型:Qwen3.8-Flash-Next究竟改变了什么

    <p>Here's a number that shouldn't make sense on first read.</p> <p>Qwen just released Qwen3.8-Flash-Next, and it activates <strong>6 billion parameters per token</strong> — while matching or beating models that activate <strong>13B (DeepSeek-V4-Flash) and 17B (Qwen3.7-Plus)</stro…

  11. dev.to — LLM tag TIER_1 English(EN) · Mariano Gobea Alcoba ·

    Qwen3.8-Flash-Next 智能、性能与价格分析!

    <h2> Architectural Evolution: A Technical Deconstruction of Qwen3.8-Flash-Next </h2> <p>The release of Qwen3.8-Flash-Next marks a significant shift in the deployment strategies for large language models (LLMs) in high-throughput, low-latency environments. As infrastructure archit…

  12. dev.to — LLM tag TIER_1 English(EN) · cz ·

    Qwen3.8-Flash-Next (2026):Qwen的Qwen4-Preview架构模型完全指南

    <h1> Qwen3.8-Flash-Next (2026): The Complete Guide to Qwen's Qwen4-Preview Architecture Model </h1> <h2> 🎯 Core Takeaways (TL;DR) </h2> <ul> <li> <strong>Qwen3.8-Flash-Next</strong> is Alibaba Qwen's open-weight preview of the architecture that will underpin <strong>Qwen4</strong…

  13. Mastodon — mastodon.social TIER_1 Polski(PL) · aisight ·

    Qwen3.8-Flash-Next 模型创AI系统成本效益新标杆,旗舰级性能每百万 tokens 仅需 0.16 美元。# si

    Model Qwen3.8-Flash-Next wyznacza nową granicę opłacalności systemów AI, oferując wydajność klasy flagowej przy cenie zaledwie 0,16 USD za milion tokenów. # si # ai # sztucznainteligencja # wiadomości # informacje # technologia https:// aisight.pl/technologia/generat ywna-ai/llm/…