PulseAugur
中
实时 22:16:40
English(EN) What are the minimum specs required to run Qwen3.8-Flash-Next?

Unsloth 加速 Qwen3.8-Flash-Next 和 GLM-5.3-Flash 性能

Unsloth 发布了更新,显著加速了 Qwen3.8-Flash-Next 和 GLM-5.3-Flash 模型的性能,提供高达 2 倍的生成速度提升和更低的 token 消耗。这些改进归功于多轮规划 (MTP) 和针对 Apple Silicon 优化的 MLX 推理等优化,能够实现更长、更快的聊天。此次发布还包括对模型加载、聊天编辑安全、本地媒体 API 和硬件兼容性(特别是 AMD GPU)的广泛增强。 AI

影响 Unsloth 的这些优化可能导致 Qwen 和 GLM 模型在各种应用中更快、更高效地部署,从而可能降低运营成本。

排序理由 对现有模型的优化库 (Unsloth) 的更新。

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 65 个来源。 我们如何撰写摘要 →

Unsloth 加速 Qwen3.8-Flash-Next 和 GLM-5.3-Flash 性能

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
对现有模型的优化库 (Unsloth) 的更新。
Source corroboration
65 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
model release, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
47 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [65]

  1. Ollama — Releases TIER_1 English(EN) · dhiltgen ·

    v0.33.1-rc0: MLX: Qwen3.8 Flash Next 支持 (#18032)

    <ul> <li> <p>MLX: Qwen3.8 Flash Next support</p> </li> <li> <p>review comments</p> </li> </ul>

  2. Unsloth — Releases TIER_1 English(EN) · danielhanchen ·

    2x 更快的 Qwen3.8-Flash + GLM-5.3-Flash MTP

    <p>Run Qwen3.8-Flash-Next and GLM-5.3-Flash up to 2x faster with MTP. MTP is enabled by default, you can still disable it.<br /> Also our new release includes 170+ training, chat, hardware, and performance improvements.</p> <h2>Highlights</h2> <ul> <li><strong>Smoother model load…

  3. Unsloth — Releases TIER_1 English(EN) · danielhanchen ·

    2x 更快的 Qwen3.8-Flash + GLM-5.3-Flash MTP

    <p>Run Qwen3.8-Flash-Next and GLM-5.3-Flash up to 2x faster with MTP. MTP is enabled by default, you can still disable it.<br /> Also our new release includes 170+ training, chat, hardware, and performance improvements.</p> <h2>Highlights</h2> <ul> <li><strong>Smoother model load…

  4. Hugging Face Trending Models TIER_1 English(EN) · nvidia ·

    nvidia/Qwen3.8-Flash-Next-NVFP4

    image-text-to-text · 1,129 downloads · 75 likes

  5. Unsloth — Releases TIER_1 English(EN) · danielhanchen ·

    Qwen3.8-Flash-Next + GLM-5.3-Flash

    <p>Qwen3.8-Flash-Next and GLM-5.3-Flash can now run locally in Unsloth!</p> <ul> <li>Run Qwen3.8-Flash-Next on 75GB RAM.</li> <li>GLM-5.3-Flash runs on 102GB of combined RAM + VRAM</li> <li>5x Faster inference if RAM offloaded</li> <li>100+ chat, reliability and performance impro…

  6. 量子位 (QbitAI) TIER_1 中文(ZH) · 量子位的朋友们 ·

    Qwen Office 首次推出 Qwen3.8-Flash,生成速度提升 100%,Token 消耗降低 75%

    8月26日晚,千问办公首发上线刚刚发布的Qwen3.8-Flash模型,同时推出标准模式。

  7. Simon Willison TIER_1 English(EN) ·

    Qwen3.8-Flash-Next

    <p><strong><a href="https://qwen.ai/blog?id=qwen3.8-flash-next">Qwen3.8-Flash-Next</a></strong></p> Another open weights model from Qwen. This one is "a multimodal MoE model that also serves as an early preview of the architecture used in Qwen4".</p> <p>It's pretty big: 125B toke…

  8. Hugging Face Trending Models TIER_1 English(EN) · RadixArk ·

    RadixArk/Qwen3.8-Flash-Next-NVFP4

    image-text-to-text · 108,962 downloads · 69 likes

  9. Hugging Face Trending Models TIER_1 Deutsch(DE) · Qwen ·

    Qwen/Qwen3.8-Flash-Next-FP8

    image-text-to-text · 451 downloads · 76 likes

  10. Hugging Face Trending Models TIER_1 Deutsch(DE) · Qwen ·

    Qwen/Qwen3.8-Flash-Next

    image-text-to-text · 2,551 downloads · 3,043 likes

  11. 雷峰网 (Leiphone) TIER_1 中文(ZH) ·

    Qwen Office 首次推出 Qwen3.8-Flash,生成速度提升 100%,Token 消耗降低 75%

    <p>8月26日晚,千问办公首发上线刚刚发布的Qwen3.8-Flash模型,同时推出标准模式。即日起,所有用户可通过全新的标准模式体验Qwen3.8-Flash。基于最新的模型,用户可以用更少的积分消耗、更快的Token吞吐速度完成任务。未来,千问办公的模型供给将只有标准和高级两种模式,95%的日常任务通过千问办公标准模式即可完成,仅5%的复杂任务需要使用高级模式。</p><p>&nbsp;</p><p style="text-align: center;"><img src="https://static.leiphone.com/uploads/n…

  12. AI Business TIER_1 English(EN) · Esther Shittu ·

    Qwen 3.8 Flash-Next 价格便宜,但存在复杂因素

    While Alibaba has kept inference and token price low, enterprises need to consider other metrics to determine if this is the right model for them.

  13. Towards AI TIER_1 English(EN) · Anubhav ·

    运行 Qwen3.8-Flash Next 无需服务器

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/you-dont-need-a-server-to-run-qwen3-8-flash-next-03e91724d196?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/2600/1*itOGf5IUED7ijnaFv8PQsg.png" width="2752…

  14. Medium — fine-tuning tag TIER_1 English(EN) · Shaaf Salman ·

    微调 Qwen3.8–27B:哪些会出错以及如何修复

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@ishaafsalman/fine-tuning-qwen3-8-27b-what-breaks-and-how-to-fix-it-d77dc46e0745?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/1361/1*nPN2TNtZ6Jin27htRqQSzQ.jpeg"…

  15. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    Show HN:在 48GB 的 Mac 上以约 12 tok/s 的速度运行 104GB Qwen3.8-Flash-Next https://github.com/carloslfu/slotstream #HackerNews #Tech #AI

    Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s https://github.com/carloslfu/slotstream # HackerNews # Tech # AI

  16. r/LocalLLaMA TIER_1 English(EN) · /u/starkruzr ·

    运行 Qwen3.8-Flash-Next 需要哪些最新的配置,才能搭配两块 3090 显卡和大量系统内存?

    <!-- SC_OFF --><div class="md"><p>we have an HPC cluster with 70 nodes, each with an A30 (roughly a 3090, same 24GB VRAM) and 1TB DDR4 RDIMM system RAM. I feel like I saw a &quot;twin 3090 and 128GB RAM&quot; recipe around here recently that could do tensor parallel and I can't f…

  17. r/LocalLLaMA TIER_1 English(EN) · /u/carteakey ·

    在 12GB 显存的显卡上本地运行 Qwen3.8-Flash-Next

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1wgiefk/running_qwen38flashnext_locally_on_a_12gb_vram/"> <img alt="Running Qwen3.8-Flash-Next locally on a 12GB VRAM card" src="https://external-preview.redd.it/zfItJPW1LJjGSCmLnlxtz-bJAvKTxm3ShNeTke9zQWU.png…

  18. r/LocalLLaMA TIER_1 English(EN) · /u/rm-rf-rm ·

    数据点:Qwen3.8-Flash-Next 在 M3 Ultra 上的 PP/TG 速度

    <!-- SC_OFF --><div class="md"><p><strong>Model File</strong>: <a href="https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF">https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF</a> Q4_K_XL</p> <p><strong>llama.cpp configuration through llama-swap:</strong></p> <pre><code>-c…

  19. r/LocalLLaMA TIER_1 English(EN) · /u/ilintar ·

    Qwen3.8 Flash Next 在 Strix Halo 上预填充速度达 1.2k t/s

    <!-- SC_OFF --><div class="md"><p>As you all know, Qwen3.8 Flash Next on mainline llama.cpp is still in a pretty experimental stage, but a lot of community forks are trying to get it to work better. There's also a closed-source solution called Halogen (<a href="https://github.com…

  20. r/LocalLLaMA TIER_1 English(EN) · /u/whiteh4cker ·

    Qwen3.8 Flash Next UD-Q4_K_XL 在 Windows 11 上使用 2x RTX 3090 以 49 tokens/s 运行 TGS。

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1wdipve/qwen38_flash_next_udq4_k_xl_49_tokenss_tgs_using/"> <img alt="Qwen3.8 Flash Next UD-Q4_K_XL 49 tokens/s TGS using 2x RTX 3090 on Windows 11." src="https://preview.redd.it/an5rtqz5nwoh1.png?width=140&am…

  21. r/LocalLLaMA TIER_1 English(EN) · /u/T_rex2700 ·

    有人似乎设法在Qwen上复制了V4.1 flash在KV上实现快速预填充的功能

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1wd4xxv/someone_apparently_managed_to_kind_of_replicate/"> <img alt="Someone apparently managed to kind of replicate what V4.1 flash does on KV for fast prefill on Qwen" src="https://preview.redd.it/9kh3f6m0at…

  22. r/LocalLLaMA TIER_1 English(EN) · /u/alfredr ·

    光速在空中:8-22 tg/s Qwen3.8-Flash-Next (Q4/Q4ish) 在 32GB M4 MacBook Air 上

    <!-- SC_OFF --><div class="md"><p>I believe I may currently hold the record for memory constrained inference for Qwen3.8–Flash-Next on Apple Silicon — needing only about 21GB of allocations.</p> <p>Introducing <a href="https://github.com/alfredr/cherenkov">Cherenkov</a>, an infer…

  23. r/LocalLLaMA TIER_1 English(EN) · /u/Extension-Bid-639 ·

    Qwen3.8-Flash-Next on 2x3090 + DDR4,第四部分:在提示运行时将专家缓存移出 GPU,预填充速度提升 2.2-2.5 倍

    <!-- SC_OFF --><div class="md"><p>Part 4 of the same box. Part 1 was 17 -&gt; 25-29 t/s with the expert cache PR, part 2 was 37-41 t/s after switching to UD-Q4_K_XL and stacking MTP on the cache, part 3 was the top-k fallback that was sorting more than it needed to. This one is a…

  24. r/LocalLLaMA TIER_1 English(EN) · /u/Beamsters ·

    Qwen3.8-Flash-Next 在 MLX-serve 上发布,支持 100 万上下文!

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1wb7p70/qwen38flashnext_on_mlxserve_1m_context_is_released/"> <img alt="Qwen3.8-Flash-Next on MLX-serve, 1m context is released!" src="https://external-preview.redd.it/amV5d2RvZ285ZW9oMcxqCVPJ9kHkkxaTFzSdDAFP7…

  25. r/LocalLLaMA TIER_1 English(EN) · /u/FantasticNature7590 ·

    Qwen3.8-Flash-Next 在 llama.cpp vs SGLang vs FreeToken 中:全上下文下首个 token 耗时 35s vs 258s。我关于引擎即将推出的新 PR 的发现。

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1waydqj/qwen38flashnext_in_llamacpp_vs_sglang_vs/"> <img alt="Qwen3.8-Flash-Next in llama.cpp vs SGLang vs FreeToken: 35s vs 258s to first token at full context. My findings on new PRs coming to engines." src=…

  26. r/LocalLLaMA TIER_1 English(EN) · /u/Extension-Bid-639 ·

    Qwen3.8-Flash-Next 在 2x3090 上:上下文长度约 119k 时解码速度提升 9-12%,已完成质量筛选

    <!-- SC_OFF --><div class="md"><p>An update to my <a href="https://inovello.dev/writeups/qwen3-flash-next-2x3090-q4-mtp/">previous post on running Flash-Next with the expert cache and MTP</a>.</p> <p>I found another useful improvement on the same dual-3090 setup: replacing the CU…

  27. r/LocalLLaMA TIER_1 English(EN) · /u/Zeeplankton ·

    您运行的是 Qwen 3.8 27b 还是 Qwen Flash Next?

    <!-- SC_OFF --><div class="md"><p>Curious about what people are preferring, if you have the hardware. I have m3 Max 96gb and both run, and largely feel identical, but prefill on qwen 27b is faster. Is there anything / anyone working on anything to improve pp with mlx?</p> <p>Bran…

  28. r/LocalLLaMA TIER_1 English(EN) · /u/DerTomsn ·

    Qwen3.8-Flash-Next-oQ4e-mtp:M4 Max 上 45 tok/s,M2 Ultra 上 25 tok/s 进行本地推理 — llm-bench.io

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w8qo79/qwen38flashnextoq4emtp_45_toks_on_m4_max_25_toks/"> <img alt="Qwen3.8-Flash-Next-oQ4e-mtp: 45 tok/s on M4 Max, 25 tok/s on M2 Ultra for local inference — llm-bench.io" src="https://external-preview.red…

  29. r/LocalLLaMA TIER_1 English(EN) · /u/HeDo88TH ·

    Qwen3.8 Flash Next - 模板对比

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w84mod/qwen38_flash_next_templates_comparison/"> <img alt="Qwen3.8 Flash Next - Templates Comparison" src="https://preview.redd.it/02geu81o8qnh1.png?width=140&amp;height=78&amp;auto=webp&amp;s=76115c95a1eaeaa…

  30. r/LocalLLaMA TIER_1 English(EN) · /u/zRevengee ·

    Qwen 3.8 Flash Next 可构建趣味游戏

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w73aak/qwen_38_flash_next_can_build_funny_games/"> <img alt="Qwen 3.8 Flash Next Can Build Funny Games" src="https://preview.redd.it/o3fnyfucwhnh1.png?width=140&amp;height=94&amp;auto=webp&amp;s=94f2f86851c6f…

  31. r/LocalLLaMA TIER_1 English(EN) · /u/memeka ·

    是只有我这么觉得,还是 Qwen3.8-Flash-Next ... 真的有很多 bug?

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w70d85/is_it_just_me_or_is_qwen38flashnext_really_buggy/"> <img alt="Is it just me or is Qwen3.8-Flash-Next ... really buggy?" src="https://preview.redd.it/l3qhlnnvahnh1.png?width=640&amp;crop=smart&amp;auto=…

  32. r/LocalLLaMA TIER_1 English(EN) · /u/BusTiny207 ·

    Qwen3.8-Flash-Next:DDR4和Tesla T4上的256k上下文,16tok/s

    <!-- SC_OFF --><div class="md"><p>I've got an refurb Dell R740 running Proxmox that I put a Tesla T4 in, mainly to run some CTC local transcription work, but thought it would be fun to try DS4 when it came out, and it was appalling at around 2 tok/s. </p> <p>However pulled it out…

  33. r/LocalLLaMA TIER_1 English(EN) · /u/Extension-Bid-639 ·

    更新:Qwen3.8-Flash-Next 在 2x3090 + DDR4 上(第二部分):25-29 -> 37-41 t/s 解码(UD-Q4_K_XL + expert cache + MTP),另附一个可供构建的分支

    <!-- SC_OFF --><div class="md"><p>This is a follow-up to my post from yesterday (17 -&gt; 25-29 t/s with the expert cache PR). Same box: 2x RTX 3090 on PCIe 3.0, dual Xeon E5-2696 v4, 188 GB DDR4-2133 LRDIMM, llama.cpp, full 261k context, f16 KV, all 48 expert layers in host RAM,…

  34. r/LocalLLaMA TIER_1 English(EN) · /u/Alternative_Will5974 ·

    Qwen3.8-Flash-Next MTP 已合并到 ik_llama.cpp (集成头或独立 -md 文件)... 5090 + 128GB 显存下 45 → 90 tok/s,在 12GB 4070 上也能运行

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w6ccgs/qwen38flashnext_mtp_merged_in_ik_llamacpp/"> <img alt="Qwen3.8-Flash-Next MTP merged in ik_llama.cpp (integrated head or separate -md file)... 45 → 90 tok/s on a 5090 + 128GB, works down to a 12GB 4070…

  35. r/LocalLLaMA TIER_1 English(EN) · /u/Extension-Bid-639 ·

    Qwen3.8-Flash-Next on 2x3090 + DDR4: 17 → 25-29 t/s 解码,得益于 expert cache PR

    <!-- SC_OFF --><div class="md"><p>Sharing some numbers because most posts on this model are either using a single 3090 or unified systems from what I've seen.</p> <p>My current setup: 2x RTX 3090 (PCIe 3.0), dual Xeon E5-2696 v4, 188 GB usable (192GB) DDR4-2133 LRDIMM, llama.cpp,…

  36. r/LocalLLaMA TIER_1 English(EN) · /u/arkham00 ·

    Qwen3.8-flash-next 处处可见腐败

    <!-- SC_OFF --><div class="md"><p>Hi, I've noticed that the model often sees &quot;garbled text&quot; in its context.</p> <p>Sometimes it declare that the tools instructions are corrupted, sometimes it is the content of some .md files, ora other files, and it freaks it out, since…

  37. r/LocalLLaMA TIER_1 English(EN) · /u/Dutchnamn ·

    Qwen3.8 Flash AP Quants

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w5ow8w/qwen38_flash_ap_quants/"> <img alt="Qwen3.8 Flash AP Quants" src="https://external-preview.redd.it/Y5IR3Y2D0kbRDs7HYpe0Hft0lcwPxlu6gvysYzKXxUE.png?width=140&amp;height=75&amp;auto=webp&amp;s=f52db6942e…

  38. r/LocalLLaMA TIER_1 English(EN) · /u/yogthos ·

    在 48GB 内存的 Mac 上以约 12 token/秒 的速度运行 104GB Qwen3.8-Flash-Next

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w4z94f/running_104gb_qwen38flashnext_on_48gb_mac_at_12/"> <img alt="Running 104GB Qwen3.8-Flash-Next on 48GB Mac at ~12 tok/s" src="https://external-preview.redd.it/XKEk4NrsCsSRJueN_-SxlfW8RQuj5_LrRWCvIiH4EPE…

  39. r/LocalLLaMA TIER_1 English(EN) · /u/vini542reddit ·

    MTP发布Qwen3.8-Flash-Next-GGUF

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w42biu/mtp_released_for_qwen38flashnextgguf/"> <img alt="MTP released for Qwen3.8-Flash-Next-GGUF" src="https://external-preview.redd.it/IwstnEKVDHtXsV0YuXUaO2VxB0-3ml6XzPZY27HuG24.png?width=640&amp;crop=smar…

  40. r/LocalLLaMA TIER_1 English(EN) · /u/FantasticNature7590 ·

    Qwen3.8-Flash-Next 在 llama.cpp 中从仅 CPU 到 96GB VRAM:8.5 到 109 token/秒,最大上下文和参数测试。我在 RTX 6000 PRO 上的发现。

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w3pl64/qwen38flashnext_in_llamacpp_from_cpuonly_to_96gb/"> <img alt="Qwen3.8-Flash-Next in llama.cpp from CPU-only to 96GB VRAM: 8.5 to 109 tok/s, max context and parameters test. My findings on RTX 6000 PRO.…

  41. r/LocalLLaMA TIER_1 English(EN) · /u/No_Algae1753 ·

    Qwen 3.8 Flash 在推理(llama.cpp)方面的当前状态如何?

    <!-- SC_OFF --><div class="md"><p>Title</p> </div><!-- SC_ON --> &#32; submitted by &#32; <a href="https://www.reddit.com/user/No_Algae1753"> /u/No_Algae1753 </a> <br /> <span><a href="https://www.reddit.com/r/LocalLLaMA/comments/1w3mneh/whats_the_current_state_of_qwen_38_flash/"…

  42. r/LocalLLaMA TIER_1 English(EN) · /u/Saren-WTAKO ·

    我的 Qwen3.8-Flash-Next 单 GB10/DGX Spark 配方,使用 Intel AutoRound int4 量化和 vLLM,fp8 n-gram 表卸载到本地 SSD 或外部 RDMA 服务器。在 mtp=3 c=1 时,代码速度约为 47.5t/s,JSON 速度约为 60t/s。Prefix cache 已开启。

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w3eser/my_qwen38flashnext_recipe_for_single_gb10dgx/"> <img alt="My Qwen3.8-Flash-Next recipe for single GB10/DGX Spark, uses Intel AutoRound int4 quant and vLLM, fp8 ngram table offloaded to local SSD or ext…

  43. r/LocalLLaMA TIER_1 English(EN) · /u/AdventurousSwim1312 ·

    Qwen3.8-Next-Flash在单块RTX 6000 Pro上速度高达240t/s

    <!-- SC_OFF --><div class="md"><p>Stumbled around a post about optimizing new Qwen up to 178t/s with a patched version of sglang : <a href="https://github.com/jpezzulli/sglang-rtxpro6000">https://github.com/jpezzulli/sglang-rtxpro6000</a></p> <p>I managed to reproduce results (ku…

  44. r/LocalLLaMA TIER_1 English(EN) · /u/Mxmtm ·

    Qwen3.8-Flash-Next 在 96GB Mac Studio 上运行(这是我的内存计算,请指出错误之处)

    <!-- SC_OFF --><div class="md"><p>Mac Studio, 96GB unified memory (<strong>M3 Ultra</strong>). I want the largest usable Qwen3.8-Flash-Next setup, and I'd rather not burn 100GB of bandwidth on the wrong download. Here's my math. Please tell me which parts are wrong.</p> <h1>What …

  45. r/LocalLLaMA TIER_1 English(EN) · /u/TemperatureOk3561 ·

    Qwen3.8-Flash-Next 的最佳 1.25 位量化

    <!-- SC_OFF --><div class="md"><p>Hello! I was looking into quantizing models and i saw how Hy4 was shrunk from 1.5 TB to 200GB with high retention in benchmarks (98% i think). I was wondering if:<br /> a) it would be worth it to attempt this method (since they had papers detaili…

  46. r/LocalLLaMA TIER_1 Deutsch(DE) · /u/trashacct383 ·

    Qwen3.8-Flash-Next-NVFP4 vs Qwen3.8-27B-FP 测试结果

    <!-- SC_OFF --><div class="md"><p>## Qwen3.8-Flash-Next-NVFP4 (inferact) vs Qwen3.8-27B-FP8 (qwen)</p> <p>Slammed with work and no time to pretty this up. Qwen wrote most of this but I checked the data.</p> <p>All tests done on the same rig, same prompts, and most tests are my re…

  47. r/LocalLLaMA TIER_1 English(EN) · /u/sloptimizer ·

    Qwen3.8-Flash-Next 将 4xR9700 变为本地 AI 强机!优化 vLLM 后单请求 120 t/s TG 和 12k t/s PP

    <!-- SC_OFF --><div class="md"><p>If you own 4xR9700 and were waiting for the model to make them shine, then I have some good news for you!</p> <p>It's running at 80-120 tokens/second for generation and 12k token/second prefill for a single request, using <a href="https://hugging…

  48. r/LocalLLaMA TIER_1 English(EN) · /u/jbro1985 ·

    在 2×RTX 3090 + 96GB DDR5 上运行 Qwen3.8-Flash-Next (125B MoE + 51B n-gram 表)。优化 Llama.cpp 和 vLLM,将专家模型卸载到 RAM,n-gram 模型卸载到 NVME (已验证并正在优化)

    <!-- SC_OFF --><div class="md"><p>Can provide configs if people are interested but did not want to do the wall of text. Below is AI assisted drafting of bullet points of what I have achieved so far:</p> <p><strong>LLAMA.CPP</strong></p> <p><strong>32 t/s decode / 463 t/s prefill …

  49. r/LocalLLaMA TIER_1 English(EN) · /u/Dutchnamn ·

    Qwen3.8 Flash Quants

    <!-- SC_OFF --><div class="md"><p>~20–30GB smaller than Unsloth/AesSedai Q4 at similar PPL. </p> <p>After several days of testing I released a set of mainline-compatible imatrix quants for Qwen3.8-Flash-Next. </p> <p>Goal: same quality band as the popular Unsloth / AesSedai Q4 bu…

  50. r/LocalLLaMA TIER_1 English(EN) · /u/TheGlobinKing ·

    Qwen3.8-27B 对比 Qwen3.8-Flash-Next 更小的量化版本?

    <!-- SC_OFF --><div class="md"><p>If you only had 128gb ram which one would be more &quot;intelligent&quot;, Qwen3.8-27B (or even 3.6) or a smaller quant of Qwen3.8-Flash-Next (Q4/Q5) ? Mostly for discussions, but also interested in coding. Thanks</p> <p>edit: I have a 128gb Halo…

  51. r/LocalLLaMA TIER_1 English(EN) · /u/betiz0 ·

    Qwen3.8-Flash-Next + MTP on Strix Halo: Vulkan 运行时笔记

    <!-- SC_OFF --><div class="md"><p>Below are the benchmark results for running Qwen3.8-Flash-Next on Strix Halo using the Vulkan backend of llama.cpp, combined with MTP model.</p> <h1>Hardware</h1> <table><thead> <tr> <th align="left">Item</th> <th align="left">Details</th> </tr> …

  52. r/LocalLLaMA TIER_1 English(EN) · /u/Acceptable_Adagio_91 ·

    在 4x3090 上运行 Qwen 3.8 Flash Next 对比 27B 是否值得?

    <!-- SC_OFF --><div class="md"><p>Can someone please tell me if it's worth running Qwen 3.8 Flash Next on 4x3090 yet over 27B?</p> <p>27B is good but damn it is indecisive. I am getting frustrated watching it get &quot;so close&quot; to solving a problem, only to do another 2 hou…

  53. r/LocalLLaMA TIER_1 English(EN) · /u/jacek2023 ·

    llama.cpp 已合并对 Qwen3.8-Flash-Next 的支持

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w03zdo/llamacpp_support_for_qwen38flashnext_has_been/"> <img alt="llama.cpp support for Qwen3.8-Flash-Next has been merged" src="https://external-preview.redd.it/clU47CA21MjdRoxLjBAdDgI-CEmGPJu2GvpJx2M9fcs.pn…

  54. dev.to — LLM tag TIER_1 English(EN) · li wujie ·

    我测试了 GLM-5.3-Flash 和 Qwen3.8-Flash 在 24 个真实任务上的表现

    <p>I test-ran both of this week's open-weight flash models against 24 small, real workloads from an actual product stack — structured extraction, SEO metadata, and code fixes — and graded everything programmatically. The short version: <strong>on quality the two models are effect…

  55. r/LocalLLaMA TIER_1 English(EN) · /u/tolitius ·

    Qwen3.8-Flash-Next:是时候更新那些基准测试了

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vzspz6/qwen38flashnext_time_to_update_those_benchmarks/"> <img alt="Qwen3.8-Flash-Next: Time to Update Those Benchmarks" src="https://preview.redd.it/6sdkwxr3swlh1.png?width=640&amp;crop=smart&amp;auto=webp&a…

  56. r/LocalLLaMA TIER_1 English(EN) · /u/StartupTim ·

    有人成功在 2x DGX Sparks 上运行 Qwen3.8 Flash 吗?

    <!-- SC_OFF --><div class="md"><p>I'm running into all sorts of errors, has anybody got Qwen3.8-Flash-Next to work on a cluster of 2x DGX Sparks?</p> <p>If so could you post your settings/recipe?</p> <p>Thanks!</p> </div><!-- SC_ON --> &#32; submitted by &#32; <a href="https://ww…

  57. r/LocalLLaMA TIER_1 English(EN) · /u/Legitimate_Hat_7852 ·

    DeepSeek V4 0731 -> Qwen 3.8 Flash -> GLM 5.3 Flash (又绕回来了!)

    <!-- SC_OFF --><div class="md"><p>Spent yesterday getting Qwen3.8 Flash and GLM 5.3 Flash up and running on my cluster of 4 x DGX Sparks with a view to replacing DeepSeek 0731... but.. really not that impressed with GLM 5.3 - overly verbose and takes for ever (was getting around …

  58. r/LocalLLaMA TIER_1 English(EN) · /u/Fancy-Snow7 ·

    运行 Qwen3.8-Flash-Next 需要哪些最低配置?

    <!-- SC_OFF --><div class="md"><p>How much system RAM? How much VRAM? How much SSD space?</p> <p>Ideally list for q3/4 but q2 might also work since I have seen 3.8 27B perform well even on q2.</p> <p>Currently I have 5070 Ti with 16GB VRAM and 48GB system RAM. I can upgrade syste…

  59. r/LocalLLaMA TIER_1 English(EN) · /u/Normal-Phone7762 ·

    Qwen3.8-Flash-Next 优于 DeepSeek V4 Pro

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vzowwo/qwen38flashnext_better_then_deepseek_v4_pro/"> <img alt="Qwen3.8-Flash-Next better then DeepSeek V4 Pro" src="https://preview.redd.it/6jri8d99xvlh1.png?width=140&amp;height=58&amp;auto=webp&amp;s=efcdf…

  60. dev.to — LLM tag TIER_1 ไทย(TH) · Nokka ·

    Qwen3.8-Flash-Next 发布,Qwen4 架构预览

    <h1> Qwen3.8-Flash-Next โมเดล 125B ที่เปิดตัวเป็น "ตัวอย่างสถาปัตยกรรม Qwen4", และทำไม Qwen ถึงครองปี 2026 </h1> <p><em>โดย Nokka (นก-กา) | 26 สิงหาคม 2026</em></p> <p><em>บทความนี้เขียนโดย AI (deepseek-v4-pro via ollama-cloud) ผ่าน Hermes Agent ภายใต้การควบคุมและตรวจสอบคุณภาพโดย…

  61. r/LocalLLaMA TIER_1 English(EN) · /u/rerri ·

    Qwen3.8-Flash-Next 明日上线

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vxwtyd/qwen38flashnext_tomorrow/"> <img alt="Qwen3.8-Flash-Next tomorrow" src="https://external-preview.redd.it/Z4u4braWhbHDqxXT7p8EHRqWEEzexwWUfFcB_ovsKUw.png?width=640&amp;crop=smart&amp;auto=webp&amp;s=2e3…

  62. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    哇!Qwen3.8-flash-next 说道 🤣 # AI

    Whoa! says Qwen3.8-flash-next 🤣 # AI

  63. Mastodon — mastodon.social TIER_1 Deutsch(DE) · aisyndicate ·

    Qwen3.8 Flash Next 登陆 DGX Spark:MTP2 测试 MTP2-speculative decoding 带来 Qwen3.8-Flash-Next 在 DGX Spark 上 15–24% 的解码吞吐量提升,占用 9 GB K

    Qwen3.8 Flash Next auf DGX Spark: MTP2 im Test MTP2-spekulative Dekodierung bringt Qwen3.8-Flash-Next auf der DGX Spark 15–24 % mehr Decode-Durchsatz bei 9 GB KV-Cache; Werkzeugqualität bleibt praktisch identisch. Der +43 %-Peak bei TG512 streut zu stark für ein belastbares Urtei…

  64. Mastodon — mastodon.social TIER_1 English(EN) · h4ckernews ·

    在 48GB 内存的 Mac 上以约 12 token/秒的速度运行 104GB Qwen3.8-Flash-Next https:// github.com/carloslfu/slotstream 评论: https:// news.ycombinator.com/item?id=4 952444

    Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s https:// github.com/carloslfu/slotstream Comments: https:// news.ycombinator.com/item?id=4 9524447 # HackerNews # Running # Qwen3 .8-Flash-Next # Mac # Performance # AI # 104GB

  65. Mastodon — mastodon.social TIER_1 Deutsch(DE) · [email protected] ·

    RT @danieltvela: 我正在 PRO 6000 上运行 Qwen3.8-Flash-Next,拥有约 56 GB DRAM。更多信息请访问 Arint.info # AI # MTP # Prefilling # Qwen3 # v

    RT @danieltvela: Ich habe Qwen3.8-Flash-Next auf einem PRO 6000 mit ~56 GB DRAM am Laufen. mehr auf Arint.info # AI # MTP # Performance # Prefilling # Qwen3 # vllm # arint_info https://x.com/danieltvela/status/2093999978473542089