PulseAugur
EN
LIVE 22:42:26

Unsloth accelerates Qwen3.8-Flash-Next and GLM-5.3-Flash performance

Unsloth has released updates that significantly accelerate the performance of Qwen3.8-Flash-Next and GLM-5.3-Flash models, offering up to 2x faster generation speeds and reduced token consumption. These improvements are attributed to optimizations like Multi-Turn Planning (MTP) and enhanced MLX inference for Apple Silicon, enabling longer and faster chats. The releases also include broader enhancements to model loading, chat editing safety, local media APIs, and hardware compatibility, particularly for AMD GPUs. AI

IMPACT These optimizations by Unsloth could lead to faster and more efficient deployment of Qwen and GLM models in various applications, potentially lowering operational costs.

RANK_REASON Updates to an optimization library (Unsloth) for existing models.

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 65 sources. How we write summaries →

Unsloth accelerates Qwen3.8-Flash-Next and GLM-5.3-Flash performance

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
Updates to an optimization library (Unsloth) for existing models.
Source corroboration
65 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
model release, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
47 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

Full methodology in our editorial standards.

COVERAGE [65]

  1. Ollama — Releases TIER_1 English(EN) · dhiltgen ·

    v0.33.1-rc0: MLX: Qwen3.8 Flash Next support (#18032)

    <ul> <li> <p>MLX: Qwen3.8 Flash Next support</p> </li> <li> <p>review comments</p> </li> </ul>

  2. Unsloth — Releases TIER_1 English(EN) · danielhanchen ·

    2x Faster Qwen3.8-Flash + GLM-5.3-Flash MTP

    <p>Run Qwen3.8-Flash-Next and GLM-5.3-Flash up to 2x faster with MTP. MTP is enabled by default, you can still disable it.<br /> Also our new release includes 170+ training, chat, hardware, and performance improvements.</p> <h2>Highlights</h2> <ul> <li><strong>Smoother model load…

  3. Unsloth — Releases TIER_1 English(EN) · danielhanchen ·

    2x Faster Qwen3.8-Flash + GLM-5.3-Flash MTP

    <p>Run Qwen3.8-Flash-Next and GLM-5.3-Flash up to 2x faster with MTP. MTP is enabled by default, you can still disable it.<br /> Also our new release includes 170+ training, chat, hardware, and performance improvements.</p> <h2>Highlights</h2> <ul> <li><strong>Smoother model load…

  4. Hugging Face Trending Models TIER_1 English(EN) · nvidia ·

    nvidia/Qwen3.8-Flash-Next-NVFP4

    image-text-to-text · 1,129 downloads · 75 likes

  5. Unsloth — Releases TIER_1 English(EN) · danielhanchen ·

    Qwen3.8-Flash-Next + GLM-5.3-Flash

    <p>Qwen3.8-Flash-Next and GLM-5.3-Flash can now run locally in Unsloth!</p> <ul> <li>Run Qwen3.8-Flash-Next on 75GB RAM.</li> <li>GLM-5.3-Flash runs on 102GB of combined RAM + VRAM</li> <li>5x Faster inference if RAM offloaded</li> <li>100+ chat, reliability and performance impro…

  6. 量子位 (QbitAI) TIER_1 中文(ZH) · 量子位的朋友们 ·

    Qwen Office Launches Qwen3.8-Flash for the First Time, Generating Speed Increased by 100%, Token Consumption Reduced by 75%

    8月26日晚,千问办公首发上线刚刚发布的Qwen3.8-Flash模型,同时推出标准模式。

  7. Simon Willison TIER_1 English(EN) ·

    Qwen3.8-Flash-Next

    <p><strong><a href="https://qwen.ai/blog?id=qwen3.8-flash-next">Qwen3.8-Flash-Next</a></strong></p> Another open weights model from Qwen. This one is "a multimodal MoE model that also serves as an early preview of the architecture used in Qwen4".</p> <p>It's pretty big: 125B toke…

  8. Hugging Face Trending Models TIER_1 English(EN) · RadixArk ·

    RadixArk/Qwen3.8-Flash-Next-NVFP4

    image-text-to-text · 108,962 downloads · 69 likes

  9. Hugging Face Trending Models TIER_1 Deutsch(DE) · Qwen ·

    Qwen/Qwen3.8-Flash-Next-FP8

    image-text-to-text · 451 downloads · 76 likes

  10. Hugging Face Trending Models TIER_1 Deutsch(DE) · Qwen ·

    Qwen/Qwen3.8-Flash-Next

    image-text-to-text · 2,551 downloads · 3,043 likes

  11. 雷峰网 (Leiphone) TIER_1 中文(ZH) ·

    Qwen Office Launches Qwen3.8-Flash for the First Time, Generating Speed Increased by 100%, Token Consumption Reduced by 75%

    <p>8月26日晚,千问办公首发上线刚刚发布的Qwen3.8-Flash模型,同时推出标准模式。即日起,所有用户可通过全新的标准模式体验Qwen3.8-Flash。基于最新的模型,用户可以用更少的积分消耗、更快的Token吞吐速度完成任务。未来,千问办公的模型供给将只有标准和高级两种模式,95%的日常任务通过千问办公标准模式即可完成,仅5%的复杂任务需要使用高级模式。</p><p>&nbsp;</p><p style="text-align: center;"><img src="https://static.leiphone.com/uploads/n…

  12. AI Business TIER_1 English(EN) · Esther Shittu ·

    Qwen 3.8 Flash-Next is Cheap, But There Are Complicating Factors

    While Alibaba has kept inference and token price low, enterprises need to consider other metrics to determine if this is the right model for them.

  13. Towards AI TIER_1 English(EN) · Anubhav ·

    You Don’t Need a Server to Run Qwen3.8-Flash Next

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/you-dont-need-a-server-to-run-qwen3-8-flash-next-03e91724d196?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/2600/1*itOGf5IUED7ijnaFv8PQsg.png" width="2752…

  14. Medium — fine-tuning tag TIER_1 English(EN) · Shaaf Salman ·

    Fine-Tuning Qwen3.8–27B: What Breaks and How to Fix It

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@ishaafsalman/fine-tuning-qwen3-8-27b-what-breaks-and-how-to-fix-it-d77dc46e0745?source=rss------fine_tuning-5"><img src="https://cdn-images-1.medium.com/max/1361/1*nPN2TNtZ6Jin27htRqQSzQ.jpeg"…

  15. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s https://github.com/carloslfu/slotstream # HackerNews # Tech # AI

    Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s https://github.com/carloslfu/slotstream # HackerNews # Tech # AI

  16. r/LocalLLaMA TIER_1 English(EN) · /u/starkruzr ·

    what's the state of the art recipe for running Qwen3.8-Flash-Next with a pair of 3090s and a ton of system RAM rn?

    <!-- SC_OFF --><div class="md"><p>we have an HPC cluster with 70 nodes, each with an A30 (roughly a 3090, same 24GB VRAM) and 1TB DDR4 RDIMM system RAM. I feel like I saw a &quot;twin 3090 and 128GB RAM&quot; recipe around here recently that could do tensor parallel and I can't f…

  17. r/LocalLLaMA TIER_1 English(EN) · /u/carteakey ·

    Running Qwen3.8-Flash-Next locally on a 12GB VRAM card

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1wgiefk/running_qwen38flashnext_locally_on_a_12gb_vram/"> <img alt="Running Qwen3.8-Flash-Next locally on a 12GB VRAM card" src="https://external-preview.redd.it/zfItJPW1LJjGSCmLnlxtz-bJAvKTxm3ShNeTke9zQWU.png…

  18. r/LocalLLaMA TIER_1 English(EN) · /u/rm-rf-rm ·

    Data point: Qwen3.8-Flash-Next PP/TG speed on M3 Ultra

    <!-- SC_OFF --><div class="md"><p><strong>Model File</strong>: <a href="https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF">https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF</a> Q4_K_XL</p> <p><strong>llama.cpp configuration through llama-swap:</strong></p> <pre><code>-c…

  19. r/LocalLLaMA TIER_1 English(EN) · /u/ilintar ·

    Qwen3.8 Flash Next now at 1.2k t/s prefill on Strix Halo

    <!-- SC_OFF --><div class="md"><p>As you all know, Qwen3.8 Flash Next on mainline llama.cpp is still in a pretty experimental stage, but a lot of community forks are trying to get it to work better. There's also a closed-source solution called Halogen (<a href="https://github.com…

  20. r/LocalLLaMA TIER_1 English(EN) · /u/whiteh4cker ·

    Qwen3.8 Flash Next UD-Q4_K_XL 49 tokens/s TGS using 2x RTX 3090 on Windows 11.

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1wdipve/qwen38_flash_next_udq4_k_xl_49_tokenss_tgs_using/"> <img alt="Qwen3.8 Flash Next UD-Q4_K_XL 49 tokens/s TGS using 2x RTX 3090 on Windows 11." src="https://preview.redd.it/an5rtqz5nwoh1.png?width=140&am…

  21. r/LocalLLaMA TIER_1 English(EN) · /u/T_rex2700 ·

    Someone apparently managed to kind of replicate what V4.1 flash does on KV for fast prefill on Qwen

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1wd4xxv/someone_apparently_managed_to_kind_of_replicate/"> <img alt="Someone apparently managed to kind of replicate what V4.1 flash does on KV for fast prefill on Qwen" src="https://preview.redd.it/9kh3f6m0at…

  22. r/LocalLLaMA TIER_1 English(EN) · /u/alfredr ·

    Faster than Light in Air: 8-22 tg/s Qwen3.8-Flash-Next (Q4/Q4ish) on a 32GB M4 MacBook Air

    <!-- SC_OFF --><div class="md"><p>I believe I may currently hold the record for memory constrained inference for Qwen3.8–Flash-Next on Apple Silicon — needing only about 21GB of allocations.</p> <p>Introducing <a href="https://github.com/alfredr/cherenkov">Cherenkov</a>, an infer…

  23. r/LocalLLaMA TIER_1 English(EN) · /u/Extension-Bid-639 ·

    Qwen3.8-Flash-Next on 2x3090 + DDR4, part 4: 2.2-2.5x faster prefill by kicking the expert cache off the GPU while the prompt runs

    <!-- SC_OFF --><div class="md"><p>Part 4 of the same box. Part 1 was 17 -&gt; 25-29 t/s with the expert cache PR, part 2 was 37-41 t/s after switching to UD-Q4_K_XL and stacking MTP on the cache, part 3 was the top-k fallback that was sorting more than it needed to. This one is a…

  24. r/LocalLLaMA TIER_1 English(EN) · /u/Beamsters ·

    Qwen3.8-Flash-Next on MLX-serve, 1m context is released!

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1wb7p70/qwen38flashnext_on_mlxserve_1m_context_is_released/"> <img alt="Qwen3.8-Flash-Next on MLX-serve, 1m context is released!" src="https://external-preview.redd.it/amV5d2RvZ285ZW9oMcxqCVPJ9kHkkxaTFzSdDAFP7…

  25. r/LocalLLaMA TIER_1 English(EN) · /u/FantasticNature7590 ·

    Qwen3.8-Flash-Next in llama.cpp vs SGLang vs FreeToken: 35s vs 258s to first token at full context. My findings on new PRs coming to engines.

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1waydqj/qwen38flashnext_in_llamacpp_vs_sglang_vs/"> <img alt="Qwen3.8-Flash-Next in llama.cpp vs SGLang vs FreeToken: 35s vs 258s to first token at full context. My findings on new PRs coming to engines." src=…

  26. r/LocalLLaMA TIER_1 English(EN) · /u/Extension-Bid-639 ·

    Qwen3.8-Flash-Next on 2x3090: 9–12% faster decode at ~119k context, with a completed quality screen

    <!-- SC_OFF --><div class="md"><p>An update to my <a href="https://inovello.dev/writeups/qwen3-flash-next-2x3090-q4-mtp/">previous post on running Flash-Next with the expert cache and MTP</a>.</p> <p>I found another useful improvement on the same dual-3090 setup: replacing the CU…

  27. r/LocalLLaMA TIER_1 English(EN) · /u/Zeeplankton ·

    Are you running Qwen 3.8 27b or Qwen Flash Next?

    <!-- SC_OFF --><div class="md"><p>Curious about what people are preferring, if you have the hardware. I have m3 Max 96gb and both run, and largely feel identical, but prefill on qwen 27b is faster. Is there anything / anyone working on anything to improve pp with mlx?</p> <p>Bran…

  28. r/LocalLLaMA TIER_1 English(EN) · /u/DerTomsn ·

    Qwen3.8-Flash-Next-oQ4e-mtp: 45 tok/s on M4 Max, 25 tok/s on M2 Ultra for local inference — llm-bench.io

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w8qo79/qwen38flashnextoq4emtp_45_toks_on_m4_max_25_toks/"> <img alt="Qwen3.8-Flash-Next-oQ4e-mtp: 45 tok/s on M4 Max, 25 tok/s on M2 Ultra for local inference — llm-bench.io" src="https://external-preview.red…

  29. r/LocalLLaMA TIER_1 English(EN) · /u/HeDo88TH ·

    Qwen3.8 Flash Next - Templates Comparison

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w84mod/qwen38_flash_next_templates_comparison/"> <img alt="Qwen3.8 Flash Next - Templates Comparison" src="https://preview.redd.it/02geu81o8qnh1.png?width=140&amp;height=78&amp;auto=webp&amp;s=76115c95a1eaeaa…

  30. r/LocalLLaMA TIER_1 English(EN) · /u/zRevengee ·

    Qwen 3.8 Flash Next Can Build Funny Games

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w73aak/qwen_38_flash_next_can_build_funny_games/"> <img alt="Qwen 3.8 Flash Next Can Build Funny Games" src="https://preview.redd.it/o3fnyfucwhnh1.png?width=140&amp;height=94&amp;auto=webp&amp;s=94f2f86851c6f…

  31. r/LocalLLaMA TIER_1 English(EN) · /u/memeka ·

    Is it just me or is Qwen3.8-Flash-Next ... really buggy?

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w70d85/is_it_just_me_or_is_qwen38flashnext_really_buggy/"> <img alt="Is it just me or is Qwen3.8-Flash-Next ... really buggy?" src="https://preview.redd.it/l3qhlnnvahnh1.png?width=640&amp;crop=smart&amp;auto=…

  32. r/LocalLLaMA TIER_1 English(EN) · /u/BusTiny207 ·

    Qwen3.8-Flash-Next: 256k context, 16tok/s on DDR4 and a Tesla T4

    <!-- SC_OFF --><div class="md"><p>I've got an refurb Dell R740 running Proxmox that I put a Tesla T4 in, mainly to run some CTC local transcription work, but thought it would be fun to try DS4 when it came out, and it was appalling at around 2 tok/s. </p> <p>However pulled it out…

  33. r/LocalLLaMA TIER_1 English(EN) · /u/Extension-Bid-639 ·

    UPDATE: Qwen3.8-Flash-Next on 2x3090 + DDR4 (Part 2): 25-29 -> 37-41 t/s decode (UD-Q4_K_XL + expert cache + MTP), plus a branch you can build

    <!-- SC_OFF --><div class="md"><p>This is a follow-up to my post from yesterday (17 -&gt; 25-29 t/s with the expert cache PR). Same box: 2x RTX 3090 on PCIe 3.0, dual Xeon E5-2696 v4, 188 GB DDR4-2133 LRDIMM, llama.cpp, full 261k context, f16 KV, all 48 expert layers in host RAM,…

  34. r/LocalLLaMA TIER_1 English(EN) · /u/Alternative_Will5974 ·

    Qwen3.8-Flash-Next MTP merged in ik_llama.cpp (integrated head or separate -md file)... 45 → 90 tok/s on a 5090 + 128GB, works down to a 12GB 4070

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w6ccgs/qwen38flashnext_mtp_merged_in_ik_llamacpp/"> <img alt="Qwen3.8-Flash-Next MTP merged in ik_llama.cpp (integrated head or separate -md file)... 45 → 90 tok/s on a 5090 + 128GB, works down to a 12GB 4070…

  35. r/LocalLLaMA TIER_1 English(EN) · /u/Extension-Bid-639 ·

    Qwen3.8-Flash-Next on 2x3090 + DDR4: 17 → 25-29 t/s decode with the expert cache PR

    <!-- SC_OFF --><div class="md"><p>Sharing some numbers because most posts on this model are either using a single 3090 or unified systems from what I've seen.</p> <p>My current setup: 2x RTX 3090 (PCIe 3.0), dual Xeon E5-2696 v4, 188 GB usable (192GB) DDR4-2133 LRDIMM, llama.cpp,…

  36. r/LocalLLaMA TIER_1 English(EN) · /u/arkham00 ·

    Qwen3.8-flash-next sees corruption everywhere

    <!-- SC_OFF --><div class="md"><p>Hi, I've noticed that the model often sees &quot;garbled text&quot; in its context.</p> <p>Sometimes it declare that the tools instructions are corrupted, sometimes it is the content of some .md files, ora other files, and it freaks it out, since…

  37. r/LocalLLaMA TIER_1 English(EN) · /u/Dutchnamn ·

    Qwen3.8 Flash AP Quants

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w5ow8w/qwen38_flash_ap_quants/"> <img alt="Qwen3.8 Flash AP Quants" src="https://external-preview.redd.it/Y5IR3Y2D0kbRDs7HYpe0Hft0lcwPxlu6gvysYzKXxUE.png?width=140&amp;height=75&amp;auto=webp&amp;s=f52db6942e…

  38. r/LocalLLaMA TIER_1 English(EN) · /u/yogthos ·

    Running 104GB Qwen3.8-Flash-Next on 48GB Mac at ~12 tok/s

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w4z94f/running_104gb_qwen38flashnext_on_48gb_mac_at_12/"> <img alt="Running 104GB Qwen3.8-Flash-Next on 48GB Mac at ~12 tok/s" src="https://external-preview.redd.it/XKEk4NrsCsSRJueN_-SxlfW8RQuj5_LrRWCvIiH4EPE…

  39. r/LocalLLaMA TIER_1 English(EN) · /u/vini542reddit ·

    MTP released for Qwen3.8-Flash-Next-GGUF

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w42biu/mtp_released_for_qwen38flashnextgguf/"> <img alt="MTP released for Qwen3.8-Flash-Next-GGUF" src="https://external-preview.redd.it/IwstnEKVDHtXsV0YuXUaO2VxB0-3ml6XzPZY27HuG24.png?width=640&amp;crop=smar…

  40. r/LocalLLaMA TIER_1 English(EN) · /u/FantasticNature7590 ·

    Qwen3.8-Flash-Next in llama.cpp from CPU-only to 96GB VRAM: 8.5 to 109 tok/s, max context and parameters test. My findings on RTX 6000 PRO.

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w3pl64/qwen38flashnext_in_llamacpp_from_cpuonly_to_96gb/"> <img alt="Qwen3.8-Flash-Next in llama.cpp from CPU-only to 96GB VRAM: 8.5 to 109 tok/s, max context and parameters test. My findings on RTX 6000 PRO.…

  41. r/LocalLLaMA TIER_1 English(EN) · /u/No_Algae1753 ·

    Whats the current state of Qwen 3.8 Flash regarding inference (llama.cpp)?

    <!-- SC_OFF --><div class="md"><p>Title</p> </div><!-- SC_ON --> &#32; submitted by &#32; <a href="https://www.reddit.com/user/No_Algae1753"> /u/No_Algae1753 </a> <br /> <span><a href="https://www.reddit.com/r/LocalLLaMA/comments/1w3mneh/whats_the_current_state_of_qwen_38_flash/"…

  42. r/LocalLLaMA TIER_1 English(EN) · /u/Saren-WTAKO ·

    My Qwen3.8-Flash-Next recipe for single GB10/DGX Spark, uses Intel AutoRound int4 quant and vLLM, fp8 ngram table offloaded to local SSD or external RDMA server. At mtp=3 c=1, code is ~47.5t/s, json is ~60t/s. Prefix cache is ON.

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w3eser/my_qwen38flashnext_recipe_for_single_gb10dgx/"> <img alt="My Qwen3.8-Flash-Next recipe for single GB10/DGX Spark, uses Intel AutoRound int4 quant and vLLM, fp8 ngram table offloaded to local SSD or ext…

  43. r/LocalLLaMA TIER_1 English(EN) · /u/AdventurousSwim1312 ·

    Qwen3.8-Next-Flash up to 240t/s on single rtx 6000 pro

    <!-- SC_OFF --><div class="md"><p>Stumbled around a post about optimizing new Qwen up to 178t/s with a patched version of sglang : <a href="https://github.com/jpezzulli/sglang-rtxpro6000">https://github.com/jpezzulli/sglang-rtxpro6000</a></p> <p>I managed to reproduce results (ku…

  44. r/LocalLLaMA TIER_1 English(EN) · /u/Mxmtm ·

    Qwen3.8-Flash-Next on a 96GB Mac Studio (here's my memory math, tell me where it's wrong)

    <!-- SC_OFF --><div class="md"><p>Mac Studio, 96GB unified memory (<strong>M3 Ultra</strong>). I want the largest usable Qwen3.8-Flash-Next setup, and I'd rather not burn 100GB of bandwidth on the wrong download. Here's my math. Please tell me which parts are wrong.</p> <h1>What …

  45. r/LocalLLaMA TIER_1 English(EN) · /u/TemperatureOk3561 ·

    Optimal 1.25 bit quantization of Qwen3.8-Flash-Next

    <!-- SC_OFF --><div class="md"><p>Hello! I was looking into quantizing models and i saw how Hy4 was shrunk from 1.5 TB to 200GB with high retention in benchmarks (98% i think). I was wondering if:<br /> a) it would be worth it to attempt this method (since they had papers detaili…

  46. r/LocalLLaMA TIER_1 Deutsch(DE) · /u/trashacct383 ·

    Qwen3.8-Flash-Next-NVFP4 vs Qwen3.8-27B-FP Test Results

    <!-- SC_OFF --><div class="md"><p>## Qwen3.8-Flash-Next-NVFP4 (inferact) vs Qwen3.8-27B-FP8 (qwen)</p> <p>Slammed with work and no time to pretty this up. Qwen wrote most of this but I checked the data.</p> <p>All tests done on the same rig, same prompts, and most tests are my re…

  47. r/LocalLLaMA TIER_1 English(EN) · /u/sloptimizer ·

    Qwen3.8-Flash-Next turns 4xR9700 into a local AI powerhouse! 120 t/s TG and 12k t/s PP single request with optimized vLLM

    <!-- SC_OFF --><div class="md"><p>If you own 4xR9700 and were waiting for the model to make them shine, then I have some good news for you!</p> <p>It's running at 80-120 tokens/second for generation and 12k token/second prefill for a single request, using <a href="https://hugging…

  48. r/LocalLLaMA TIER_1 English(EN) · /u/jbro1985 ·

    Running Qwen3.8-Flash-Next (125B MoE + 51B n-gram table) on 2×RTX 3090 + 96GB DDR5. Optimised Llama.cpp and vLLM with experts offload to RAM and n-gram offload to NVME (proven and being optimised)

    <!-- SC_OFF --><div class="md"><p>Can provide configs if people are interested but did not want to do the wall of text. Below is AI assisted drafting of bullet points of what I have achieved so far:</p> <p><strong>LLAMA.CPP</strong></p> <p><strong>32 t/s decode / 463 t/s prefill …

  49. r/LocalLLaMA TIER_1 English(EN) · /u/Dutchnamn ·

    Qwen3.8 Flash Quants

    <!-- SC_OFF --><div class="md"><p>~20–30GB smaller than Unsloth/AesSedai Q4 at similar PPL. </p> <p>After several days of testing I released a set of mainline-compatible imatrix quants for Qwen3.8-Flash-Next. </p> <p>Goal: same quality band as the popular Unsloth / AesSedai Q4 bu…

  50. r/LocalLLaMA TIER_1 English(EN) · /u/TheGlobinKing ·

    Qwen3.8-27B vs Qwen3.8-Flash-Next smaller quant?

    <!-- SC_OFF --><div class="md"><p>If you only had 128gb ram which one would be more &quot;intelligent&quot;, Qwen3.8-27B (or even 3.6) or a smaller quant of Qwen3.8-Flash-Next (Q4/Q5) ? Mostly for discussions, but also interested in coding. Thanks</p> <p>edit: I have a 128gb Halo…

  51. r/LocalLLaMA TIER_1 English(EN) · /u/betiz0 ·

    Qwen3.8-Flash-Next + MTP on Strix Halo: Vulkan Runtime Notes

    <!-- SC_OFF --><div class="md"><p>Below are the benchmark results for running Qwen3.8-Flash-Next on Strix Halo using the Vulkan backend of llama.cpp, combined with MTP model.</p> <h1>Hardware</h1> <table><thead> <tr> <th align="left">Item</th> <th align="left">Details</th> </tr> …

  52. r/LocalLLaMA TIER_1 English(EN) · /u/Acceptable_Adagio_91 ·

    Is it worth running Qwen 3.8 Flash Next on 4x3090 vs 27B?

    <!-- SC_OFF --><div class="md"><p>Can someone please tell me if it's worth running Qwen 3.8 Flash Next on 4x3090 yet over 27B?</p> <p>27B is good but damn it is indecisive. I am getting frustrated watching it get &quot;so close&quot; to solving a problem, only to do another 2 hou…

  53. r/LocalLLaMA TIER_1 English(EN) · /u/jacek2023 ·

    llama.cpp support for Qwen3.8-Flash-Next has been merged

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w03zdo/llamacpp_support_for_qwen38flashnext_has_been/"> <img alt="llama.cpp support for Qwen3.8-Flash-Next has been merged" src="https://external-preview.redd.it/clU47CA21MjdRoxLjBAdDgI-CEmGPJu2GvpJx2M9fcs.pn…

  54. dev.to — LLM tag TIER_1 English(EN) · li wujie ·

    I Tested GLM-5.3-Flash and Qwen3.8-Flash on 24 Real Tasks

    <p>I test-ran both of this week's open-weight flash models against 24 small, real workloads from an actual product stack — structured extraction, SEO metadata, and code fixes — and graded everything programmatically. The short version: <strong>on quality the two models are effect…

  55. r/LocalLLaMA TIER_1 English(EN) · /u/tolitius ·

    Qwen3.8-Flash-Next: Time to Update Those Benchmarks

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vzspz6/qwen38flashnext_time_to_update_those_benchmarks/"> <img alt="Qwen3.8-Flash-Next: Time to Update Those Benchmarks" src="https://preview.redd.it/6sdkwxr3swlh1.png?width=640&amp;crop=smart&amp;auto=webp&a…

  56. r/LocalLLaMA TIER_1 English(EN) · /u/StartupTim ·

    Has anybody got Qwen3.8 Flash to work on 2x DGX Sparks?

    <!-- SC_OFF --><div class="md"><p>I'm running into all sorts of errors, has anybody got Qwen3.8-Flash-Next to work on a cluster of 2x DGX Sparks?</p> <p>If so could you post your settings/recipe?</p> <p>Thanks!</p> </div><!-- SC_ON --> &#32; submitted by &#32; <a href="https://ww…

  57. r/LocalLLaMA TIER_1 English(EN) · /u/Legitimate_Hat_7852 ·

    DeepSeek V4 0731 -> Qwen 3.8 Flash -> GLM 5.3 Flash (and back again!)

    <!-- SC_OFF --><div class="md"><p>Spent yesterday getting Qwen3.8 Flash and GLM 5.3 Flash up and running on my cluster of 4 x DGX Sparks with a view to replacing DeepSeek 0731... but.. really not that impressed with GLM 5.3 - overly verbose and takes for ever (was getting around …

  58. r/LocalLLaMA TIER_1 English(EN) · /u/Fancy-Snow7 ·

    What are the minimum specs required to run Qwen3.8-Flash-Next?

    <!-- SC_OFF --><div class="md"><p>How much system RAM? How much VRAM? How much SSD space?</p> <p>Ideally list for q3/4 but q2 might also work since I have seen 3.8 27B perform well even on q2.</p> <p>Currently I have 5070 Ti with 16GB VRAM and 48GB system RAM. I can upgrade syste…

  59. r/LocalLLaMA TIER_1 English(EN) · /u/Normal-Phone7762 ·

    Qwen3.8-Flash-Next better then DeepSeek V4 Pro

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vzowwo/qwen38flashnext_better_then_deepseek_v4_pro/"> <img alt="Qwen3.8-Flash-Next better then DeepSeek V4 Pro" src="https://preview.redd.it/6jri8d99xvlh1.png?width=140&amp;height=58&amp;auto=webp&amp;s=efcdf…

  60. dev.to — LLM tag TIER_1 ไทย(TH) · Nokka ·

    Qwen3.8-Flash-Next Launched, Qwen4 Architecture Preview

    <h1> Qwen3.8-Flash-Next โมเดล 125B ที่เปิดตัวเป็น "ตัวอย่างสถาปัตยกรรม Qwen4", และทำไม Qwen ถึงครองปี 2026 </h1> <p><em>โดย Nokka (นก-กา) | 26 สิงหาคม 2026</em></p> <p><em>บทความนี้เขียนโดย AI (deepseek-v4-pro via ollama-cloud) ผ่าน Hermes Agent ภายใต้การควบคุมและตรวจสอบคุณภาพโดย…

  61. r/LocalLLaMA TIER_1 English(EN) · /u/rerri ·

    Qwen3.8-Flash-Next tomorrow

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vxwtyd/qwen38flashnext_tomorrow/"> <img alt="Qwen3.8-Flash-Next tomorrow" src="https://external-preview.redd.it/Z4u4braWhbHDqxXT7p8EHRqWEEzexwWUfFcB_ovsKUw.png?width=640&amp;crop=smart&amp;auto=webp&amp;s=2e3…

  62. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    Whoa! says Qwen3.8-flash-next 🤣 # AI

    Whoa! says Qwen3.8-flash-next 🤣 # AI

  63. Mastodon — mastodon.social TIER_1 Deutsch(DE) · aisyndicate ·

    Qwen3.8 Flash Next on DGX Spark: MTP2 in Test MTP2-speculative decoding brings Qwen3.8-Flash-Next on DGX Spark 15–24% more decode throughput at 9 GB K

    Qwen3.8 Flash Next auf DGX Spark: MTP2 im Test MTP2-spekulative Dekodierung bringt Qwen3.8-Flash-Next auf der DGX Spark 15–24 % mehr Decode-Durchsatz bei 9 GB KV-Cache; Werkzeugqualität bleibt praktisch identisch. Der +43 %-Peak bei TG512 streut zu stark für ein belastbares Urtei…

  64. Mastodon — mastodon.social TIER_1 English(EN) · h4ckernews ·

    Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s https:// github.com/carloslfu/slotstream Comments: https:// news.ycombinator.com/item?id=4 952444

    Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s https:// github.com/carloslfu/slotstream Comments: https:// news.ycombinator.com/item?id=4 9524447 # HackerNews # Running # Qwen3 .8-Flash-Next # Mac # Performance # AI # 104GB

  65. Mastodon — mastodon.social TIER_1 Deutsch(DE) · [email protected] ·

    RT @danieltvela: I am running Qwen3.8-Flash-Next on a PRO 6000 with ~56 GB DRAM. More on Arint.info # AI # MTP # Prefilling # Qwen3 # v

    RT @danieltvela: Ich habe Qwen3.8-Flash-Next auf einem PRO 6000 mit ~56 GB DRAM am Laufen. mehr auf Arint.info # AI # MTP # Performance # Prefilling # Qwen3 # vllm # arint_info https://x.com/danieltvela/status/2093999978473542089