Unsloth accelerates Qwen3.8-Flash-Next and GLM-5.3-Flash performance
ByPulseAugur Editorial·[65 sources]·
Unsloth has released updates that significantly accelerate the performance of Qwen3.8-Flash-Next and GLM-5.3-Flash models, offering up to 2x faster generation speeds and reduced token consumption. These improvements are attributed to optimizations like Multi-Turn Planning (MTP) and enhanced MLX inference for Apple Silicon, enabling longer and faster chats. The releases also include broader enhancements to model loading, chat editing safety, local media APIs, and hardware compatibility, particularly for AMD GPUs.
AI
IMPACT
These optimizations by Unsloth could lead to faster and more efficient deployment of Qwen and GLM models in various applications, potentially lowering operational costs.
RANK_REASON
Updates to an optimization library (Unsloth) for existing models.
<p>Run Qwen3.8-Flash-Next and GLM-5.3-Flash up to 2x faster with MTP. MTP is enabled by default, you can still disable it.<br /> Also our new release includes 170+ training, chat, hardware, and performance improvements.</p> <h2>Highlights</h2> <ul> <li><strong>Smoother model load…
<p>Run Qwen3.8-Flash-Next and GLM-5.3-Flash up to 2x faster with MTP. MTP is enabled by default, you can still disable it.<br /> Also our new release includes 170+ training, chat, hardware, and performance improvements.</p> <h2>Highlights</h2> <ul> <li><strong>Smoother model load…
Hugging Face Trending Models
TIER_1English(EN)·nvidia·
<p>Qwen3.8-Flash-Next and GLM-5.3-Flash can now run locally in Unsloth!</p> <ul> <li>Run Qwen3.8-Flash-Next on 75GB RAM.</li> <li>GLM-5.3-Flash runs on 102GB of combined RAM + VRAM</li> <li>5x Faster inference if RAM offloaded</li> <li>100+ chat, reliability and performance impro…
<p><strong><a href="https://qwen.ai/blog?id=qwen3.8-flash-next">Qwen3.8-Flash-Next</a></strong></p> Another open weights model from Qwen. This one is "a multimodal MoE model that also serves as an early preview of the architecture used in Qwen4".</p> <p>It's pretty big: 125B toke…
Hugging Face Trending Models
TIER_1English(EN)·RadixArk·
<!-- SC_OFF --><div class="md"><p>we have an HPC cluster with 70 nodes, each with an A30 (roughly a 3090, same 24GB VRAM) and 1TB DDR4 RDIMM system RAM. I feel like I saw a "twin 3090 and 128GB RAM" recipe around here recently that could do tensor parallel and I can't f…
<!-- SC_OFF --><div class="md"><p>As you all know, Qwen3.8 Flash Next on mainline llama.cpp is still in a pretty experimental stage, but a lot of community forks are trying to get it to work better. There's also a closed-source solution called Halogen (<a href="https://github.com…
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1wdipve/qwen38_flash_next_udq4_k_xl_49_tokenss_tgs_using/"> <img alt="Qwen3.8 Flash Next UD-Q4_K_XL 49 tokens/s TGS using 2x RTX 3090 on Windows 11." src="https://preview.redd.it/an5rtqz5nwoh1.png?width=140&am…
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1wd4xxv/someone_apparently_managed_to_kind_of_replicate/"> <img alt="Someone apparently managed to kind of replicate what V4.1 flash does on KV for fast prefill on Qwen" src="https://preview.redd.it/9kh3f6m0at…
<!-- SC_OFF --><div class="md"><p>I believe I may currently hold the record for memory constrained inference for Qwen3.8–Flash-Next on Apple Silicon — needing only about 21GB of allocations.</p> <p>Introducing <a href="https://github.com/alfredr/cherenkov">Cherenkov</a>, an infer…
<!-- SC_OFF --><div class="md"><p>Part 4 of the same box. Part 1 was 17 -> 25-29 t/s with the expert cache PR, part 2 was 37-41 t/s after switching to UD-Q4_K_XL and stacking MTP on the cache, part 3 was the top-k fallback that was sorting more than it needed to. This one is a…
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1waydqj/qwen38flashnext_in_llamacpp_vs_sglang_vs/"> <img alt="Qwen3.8-Flash-Next in llama.cpp vs SGLang vs FreeToken: 35s vs 258s to first token at full context. My findings on new PRs coming to engines." src=…
<!-- SC_OFF --><div class="md"><p>An update to my <a href="https://inovello.dev/writeups/qwen3-flash-next-2x3090-q4-mtp/">previous post on running Flash-Next with the expert cache and MTP</a>.</p> <p>I found another useful improvement on the same dual-3090 setup: replacing the CU…
<!-- SC_OFF --><div class="md"><p>Curious about what people are preferring, if you have the hardware. I have m3 Max 96gb and both run, and largely feel identical, but prefill on qwen 27b is faster. Is there anything / anyone working on anything to improve pp with mlx?</p> <p>Bran…
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w70d85/is_it_just_me_or_is_qwen38flashnext_really_buggy/"> <img alt="Is it just me or is Qwen3.8-Flash-Next ... really buggy?" src="https://preview.redd.it/l3qhlnnvahnh1.png?width=640&crop=smart&auto=…
<!-- SC_OFF --><div class="md"><p>I've got an refurb Dell R740 running Proxmox that I put a Tesla T4 in, mainly to run some CTC local transcription work, but thought it would be fun to try DS4 when it came out, and it was appalling at around 2 tok/s. </p> <p>However pulled it out…
<!-- SC_OFF --><div class="md"><p>This is a follow-up to my post from yesterday (17 -> 25-29 t/s with the expert cache PR). Same box: 2x RTX 3090 on PCIe 3.0, dual Xeon E5-2696 v4, 188 GB DDR4-2133 LRDIMM, llama.cpp, full 261k context, f16 KV, all 48 expert layers in host RAM,…
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w6ccgs/qwen38flashnext_mtp_merged_in_ik_llamacpp/"> <img alt="Qwen3.8-Flash-Next MTP merged in ik_llama.cpp (integrated head or separate -md file)... 45 → 90 tok/s on a 5090 + 128GB, works down to a 12GB 4070…
<!-- SC_OFF --><div class="md"><p>Sharing some numbers because most posts on this model are either using a single 3090 or unified systems from what I've seen.</p> <p>My current setup: 2x RTX 3090 (PCIe 3.0), dual Xeon E5-2696 v4, 188 GB usable (192GB) DDR4-2133 LRDIMM, llama.cpp,…
<!-- SC_OFF --><div class="md"><p>Hi, I've noticed that the model often sees "garbled text" in its context.</p> <p>Sometimes it declare that the tools instructions are corrupted, sometimes it is the content of some .md files, ora other files, and it freaks it out, since…
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w4z94f/running_104gb_qwen38flashnext_on_48gb_mac_at_12/"> <img alt="Running 104GB Qwen3.8-Flash-Next on 48GB Mac at ~12 tok/s" src="https://external-preview.redd.it/XKEk4NrsCsSRJueN_-SxlfW8RQuj5_LrRWCvIiH4EPE…
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w42biu/mtp_released_for_qwen38flashnextgguf/"> <img alt="MTP released for Qwen3.8-Flash-Next-GGUF" src="https://external-preview.redd.it/IwstnEKVDHtXsV0YuXUaO2VxB0-3ml6XzPZY27HuG24.png?width=640&crop=smar…
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w3pl64/qwen38flashnext_in_llamacpp_from_cpuonly_to_96gb/"> <img alt="Qwen3.8-Flash-Next in llama.cpp from CPU-only to 96GB VRAM: 8.5 to 109 tok/s, max context and parameters test. My findings on RTX 6000 PRO.…
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w3eser/my_qwen38flashnext_recipe_for_single_gb10dgx/"> <img alt="My Qwen3.8-Flash-Next recipe for single GB10/DGX Spark, uses Intel AutoRound int4 quant and vLLM, fp8 ngram table offloaded to local SSD or ext…
<!-- SC_OFF --><div class="md"><p>Stumbled around a post about optimizing new Qwen up to 178t/s with a patched version of sglang : <a href="https://github.com/jpezzulli/sglang-rtxpro6000">https://github.com/jpezzulli/sglang-rtxpro6000</a></p> <p>I managed to reproduce results (ku…
<!-- SC_OFF --><div class="md"><p>Mac Studio, 96GB unified memory (<strong>M3 Ultra</strong>). I want the largest usable Qwen3.8-Flash-Next setup, and I'd rather not burn 100GB of bandwidth on the wrong download. Here's my math. Please tell me which parts are wrong.</p> <h1>What …
<!-- SC_OFF --><div class="md"><p>Hello! I was looking into quantizing models and i saw how Hy4 was shrunk from 1.5 TB to 200GB with high retention in benchmarks (98% i think). I was wondering if:<br /> a) it would be worth it to attempt this method (since they had papers detaili…
<!-- SC_OFF --><div class="md"><p>## Qwen3.8-Flash-Next-NVFP4 (inferact) vs Qwen3.8-27B-FP8 (qwen)</p> <p>Slammed with work and no time to pretty this up. Qwen wrote most of this but I checked the data.</p> <p>All tests done on the same rig, same prompts, and most tests are my re…
<!-- SC_OFF --><div class="md"><p>If you own 4xR9700 and were waiting for the model to make them shine, then I have some good news for you!</p> <p>It's running at 80-120 tokens/second for generation and 12k token/second prefill for a single request, using <a href="https://hugging…
<!-- SC_OFF --><div class="md"><p>Can provide configs if people are interested but did not want to do the wall of text. Below is AI assisted drafting of bullet points of what I have achieved so far:</p> <p><strong>LLAMA.CPP</strong></p> <p><strong>32 t/s decode / 463 t/s prefill …
<!-- SC_OFF --><div class="md"><p>~20–30GB smaller than Unsloth/AesSedai Q4 at similar PPL. </p> <p>After several days of testing I released a set of mainline-compatible imatrix quants for Qwen3.8-Flash-Next. </p> <p>Goal: same quality band as the popular Unsloth / AesSedai Q4 bu…
<!-- SC_OFF --><div class="md"><p>If you only had 128gb ram which one would be more "intelligent", Qwen3.8-27B (or even 3.6) or a smaller quant of Qwen3.8-Flash-Next (Q4/Q5) ? Mostly for discussions, but also interested in coding. Thanks</p> <p>edit: I have a 128gb Halo…
<!-- SC_OFF --><div class="md"><p>Below are the benchmark results for running Qwen3.8-Flash-Next on Strix Halo using the Vulkan backend of llama.cpp, combined with MTP model.</p> <h1>Hardware</h1> <table><thead> <tr> <th align="left">Item</th> <th align="left">Details</th> </tr> …
<!-- SC_OFF --><div class="md"><p>Can someone please tell me if it's worth running Qwen 3.8 Flash Next on 4x3090 yet over 27B?</p> <p>27B is good but damn it is indecisive. I am getting frustrated watching it get "so close" to solving a problem, only to do another 2 hou…
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w03zdo/llamacpp_support_for_qwen38flashnext_has_been/"> <img alt="llama.cpp support for Qwen3.8-Flash-Next has been merged" src="https://external-preview.redd.it/clU47CA21MjdRoxLjBAdDgI-CEmGPJu2GvpJx2M9fcs.pn…
<p>I test-ran both of this week's open-weight flash models against 24 small, real workloads from an actual product stack — structured extraction, SEO metadata, and code fixes — and graded everything programmatically. The short version: <strong>on quality the two models are effect…
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vzspz6/qwen38flashnext_time_to_update_those_benchmarks/"> <img alt="Qwen3.8-Flash-Next: Time to Update Those Benchmarks" src="https://preview.redd.it/6sdkwxr3swlh1.png?width=640&crop=smart&auto=webp&a…
<!-- SC_OFF --><div class="md"><p>I'm running into all sorts of errors, has anybody got Qwen3.8-Flash-Next to work on a cluster of 2x DGX Sparks?</p> <p>If so could you post your settings/recipe?</p> <p>Thanks!</p> </div><!-- SC_ON -->   submitted by   <a href="https://ww…
<!-- SC_OFF --><div class="md"><p>Spent yesterday getting Qwen3.8 Flash and GLM 5.3 Flash up and running on my cluster of 4 x DGX Sparks with a view to replacing DeepSeek 0731... but.. really not that impressed with GLM 5.3 - overly verbose and takes for ever (was getting around …
<!-- SC_OFF --><div class="md"><p>How much system RAM? How much VRAM? How much SSD space?</p> <p>Ideally list for q3/4 but q2 might also work since I have seen 3.8 27B perform well even on q2.</p> <p>Currently I have 5070 Ti with 16GB VRAM and 48GB system RAM. I can upgrade syste…
Qwen3.8 Flash Next auf DGX Spark: MTP2 im Test MTP2-spekulative Dekodierung bringt Qwen3.8-Flash-Next auf der DGX Spark 15–24 % mehr Decode-Durchsatz bei 9 GB KV-Cache; Werkzeugqualität bleibt praktisch identisch. Der +43 %-Peak bei TG512 streut zu stark für ein belastbares Urtei…
RT @danieltvela: Ich habe Qwen3.8-Flash-Next auf einem PRO 6000 mit ~56 GB DRAM am Laufen. mehr auf Arint.info # AI # MTP # Performance # Prefilling # Qwen3 # vllm # arint_info https://x.com/danieltvela/status/2093999978473542089