<p>Run Qwen3.8-Flash-Next and GLM-5.3-Flash up to 2x faster with MTP. MTP is enabled by default, you can still disable it.<br /> Also our new release includes 170+ training, chat, hardware, and performance improvements.</p> <h2>Highlights</h2> <ul> <li><strong>Smoother model load…
<p>Run Qwen3.8-Flash-Next and GLM-5.3-Flash up to 2x faster with MTP. MTP is enabled by default, you can still disable it.<br /> Also our new release includes 170+ training, chat, hardware, and performance improvements.</p> <h2>Highlights</h2> <ul> <li><strong>Smoother model load…
Hugging Face Trending Models
TIER_1English(EN)·nvidia·
<p>Qwen3.8-Flash-Next and GLM-5.3-Flash can now run locally in Unsloth!</p> <ul> <li>Run Qwen3.8-Flash-Next on 75GB RAM.</li> <li>GLM-5.3-Flash runs on 102GB of combined RAM + VRAM</li> <li>5x Faster inference if RAM offloaded</li> <li>100+ chat, reliability and performance impro…
<p><strong><a href="https://qwen.ai/blog?id=qwen3.8-flash-next">Qwen3.8-Flash-Next</a></strong></p> Another open weights model from Qwen. This one is "a multimodal MoE model that also serves as an early preview of the architecture used in Qwen4".</p> <p>It's pretty big: 125B toke…
Hugging Face Trending Models
TIER_1English(EN)·RadixArk·
<!-- SC_OFF --><div class="md"><p>we have an HPC cluster with 70 nodes, each with an A30 (roughly a 3090, same 24GB VRAM) and 1TB DDR4 RDIMM system RAM. I feel like I saw a "twin 3090 and 128GB RAM" recipe around here recently that could do tensor parallel and I can't f…
<!-- SC_OFF --><div class="md"><p>As you all know, Qwen3.8 Flash Next on mainline llama.cpp is still in a pretty experimental stage, but a lot of community forks are trying to get it to work better. There's also a closed-source solution called Halogen (<a href="https://github.com…
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1wdipve/qwen38_flash_next_udq4_k_xl_49_tokenss_tgs_using/"> <img alt="Qwen3.8 Flash Next UD-Q4_K_XL 49 tokens/s TGS using 2x RTX 3090 on Windows 11." src="https://preview.redd.it/an5rtqz5nwoh1.png?width=140&am…
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1wd4xxv/someone_apparently_managed_to_kind_of_replicate/"> <img alt="Someone apparently managed to kind of replicate what V4.1 flash does on KV for fast prefill on Qwen" src="https://preview.redd.it/9kh3f6m0at…
<!-- SC_OFF --><div class="md"><p>I believe I may currently hold the record for memory constrained inference for Qwen3.8–Flash-Next on Apple Silicon — needing only about 21GB of allocations.</p> <p>Introducing <a href="https://github.com/alfredr/cherenkov">Cherenkov</a>, an infer…
<!-- SC_OFF --><div class="md"><p>Part 4 of the same box. Part 1 was 17 -> 25-29 t/s with the expert cache PR, part 2 was 37-41 t/s after switching to UD-Q4_K_XL and stacking MTP on the cache, part 3 was the top-k fallback that was sorting more than it needed to. This one is a…
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1waydqj/qwen38flashnext_in_llamacpp_vs_sglang_vs/"> <img alt="Qwen3.8-Flash-Next in llama.cpp vs SGLang vs FreeToken: 35s vs 258s to first token at full context. My findings on new PRs coming to engines." src=…
<!-- SC_OFF --><div class="md"><p>An update to my <a href="https://inovello.dev/writeups/qwen3-flash-next-2x3090-q4-mtp/">previous post on running Flash-Next with the expert cache and MTP</a>.</p> <p>I found another useful improvement on the same dual-3090 setup: replacing the CU…
<!-- SC_OFF --><div class="md"><p>Curious about what people are preferring, if you have the hardware. I have m3 Max 96gb and both run, and largely feel identical, but prefill on qwen 27b is faster. Is there anything / anyone working on anything to improve pp with mlx?</p> <p>Bran…
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w70d85/is_it_just_me_or_is_qwen38flashnext_really_buggy/"> <img alt="Is it just me or is Qwen3.8-Flash-Next ... really buggy?" src="https://preview.redd.it/l3qhlnnvahnh1.png?width=640&crop=smart&auto=…
<!-- SC_OFF --><div class="md"><p>I've got an refurb Dell R740 running Proxmox that I put a Tesla T4 in, mainly to run some CTC local transcription work, but thought it would be fun to try DS4 when it came out, and it was appalling at around 2 tok/s. </p> <p>However pulled it out…
<!-- SC_OFF --><div class="md"><p>This is a follow-up to my post from yesterday (17 -> 25-29 t/s with the expert cache PR). Same box: 2x RTX 3090 on PCIe 3.0, dual Xeon E5-2696 v4, 188 GB DDR4-2133 LRDIMM, llama.cpp, full 261k context, f16 KV, all 48 expert layers in host RAM,…
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w6ccgs/qwen38flashnext_mtp_merged_in_ik_llamacpp/"> <img alt="Qwen3.8-Flash-Next MTP merged in ik_llama.cpp (integrated head or separate -md file)... 45 → 90 tok/s on a 5090 + 128GB, works down to a 12GB 4070…
<!-- SC_OFF --><div class="md"><p>Sharing some numbers because most posts on this model are either using a single 3090 or unified systems from what I've seen.</p> <p>My current setup: 2x RTX 3090 (PCIe 3.0), dual Xeon E5-2696 v4, 188 GB usable (192GB) DDR4-2133 LRDIMM, llama.cpp,…
<!-- SC_OFF --><div class="md"><p>Hi, I've noticed that the model often sees "garbled text" in its context.</p> <p>Sometimes it declare that the tools instructions are corrupted, sometimes it is the content of some .md files, ora other files, and it freaks it out, since…
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w4z94f/running_104gb_qwen38flashnext_on_48gb_mac_at_12/"> <img alt="Running 104GB Qwen3.8-Flash-Next on 48GB Mac at ~12 tok/s" src="https://external-preview.redd.it/XKEk4NrsCsSRJueN_-SxlfW8RQuj5_LrRWCvIiH4EPE…
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w42biu/mtp_released_for_qwen38flashnextgguf/"> <img alt="MTP released for Qwen3.8-Flash-Next-GGUF" src="https://external-preview.redd.it/IwstnEKVDHtXsV0YuXUaO2VxB0-3ml6XzPZY27HuG24.png?width=640&crop=smar…
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w3pl64/qwen38flashnext_in_llamacpp_from_cpuonly_to_96gb/"> <img alt="Qwen3.8-Flash-Next in llama.cpp from CPU-only to 96GB VRAM: 8.5 to 109 tok/s, max context and parameters test. My findings on RTX 6000 PRO.…
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w3eser/my_qwen38flashnext_recipe_for_single_gb10dgx/"> <img alt="My Qwen3.8-Flash-Next recipe for single GB10/DGX Spark, uses Intel AutoRound int4 quant and vLLM, fp8 ngram table offloaded to local SSD or ext…
<!-- SC_OFF --><div class="md"><p>Stumbled around a post about optimizing new Qwen up to 178t/s with a patched version of sglang : <a href="https://github.com/jpezzulli/sglang-rtxpro6000">https://github.com/jpezzulli/sglang-rtxpro6000</a></p> <p>I managed to reproduce results (ku…
<!-- SC_OFF --><div class="md"><p>Mac Studio, 96GB unified memory (<strong>M3 Ultra</strong>). I want the largest usable Qwen3.8-Flash-Next setup, and I'd rather not burn 100GB of bandwidth on the wrong download. Here's my math. Please tell me which parts are wrong.</p> <h1>What …
<!-- SC_OFF --><div class="md"><p>Hello! I was looking into quantizing models and i saw how Hy4 was shrunk from 1.5 TB to 200GB with high retention in benchmarks (98% i think). I was wondering if:<br /> a) it would be worth it to attempt this method (since they had papers detaili…
<!-- SC_OFF --><div class="md"><p>## Qwen3.8-Flash-Next-NVFP4 (inferact) vs Qwen3.8-27B-FP8 (qwen)</p> <p>Slammed with work and no time to pretty this up. Qwen wrote most of this but I checked the data.</p> <p>All tests done on the same rig, same prompts, and most tests are my re…
<!-- SC_OFF --><div class="md"><p>If you own 4xR9700 and were waiting for the model to make them shine, then I have some good news for you!</p> <p>It's running at 80-120 tokens/second for generation and 12k token/second prefill for a single request, using <a href="https://hugging…
<!-- SC_OFF --><div class="md"><p>Can provide configs if people are interested but did not want to do the wall of text. Below is AI assisted drafting of bullet points of what I have achieved so far:</p> <p><strong>LLAMA.CPP</strong></p> <p><strong>32 t/s decode / 463 t/s prefill …
<!-- SC_OFF --><div class="md"><p>~20–30GB smaller than Unsloth/AesSedai Q4 at similar PPL. </p> <p>After several days of testing I released a set of mainline-compatible imatrix quants for Qwen3.8-Flash-Next. </p> <p>Goal: same quality band as the popular Unsloth / AesSedai Q4 bu…
<!-- SC_OFF --><div class="md"><p>If you only had 128gb ram which one would be more "intelligent", Qwen3.8-27B (or even 3.6) or a smaller quant of Qwen3.8-Flash-Next (Q4/Q5) ? Mostly for discussions, but also interested in coding. Thanks</p> <p>edit: I have a 128gb Halo…
<!-- SC_OFF --><div class="md"><p>Below are the benchmark results for running Qwen3.8-Flash-Next on Strix Halo using the Vulkan backend of llama.cpp, combined with MTP model.</p> <h1>Hardware</h1> <table><thead> <tr> <th align="left">Item</th> <th align="left">Details</th> </tr> …
<!-- SC_OFF --><div class="md"><p>Can someone please tell me if it's worth running Qwen 3.8 Flash Next on 4x3090 yet over 27B?</p> <p>27B is good but damn it is indecisive. I am getting frustrated watching it get "so close" to solving a problem, only to do another 2 hou…
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w03zdo/llamacpp_support_for_qwen38flashnext_has_been/"> <img alt="llama.cpp support for Qwen3.8-Flash-Next has been merged" src="https://external-preview.redd.it/clU47CA21MjdRoxLjBAdDgI-CEmGPJu2GvpJx2M9fcs.pn…
<p>I test-ran both of this week's open-weight flash models against 24 small, real workloads from an actual product stack — structured extraction, SEO metadata, and code fixes — and graded everything programmatically. The short version: <strong>on quality the two models are effect…
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vzspz6/qwen38flashnext_time_to_update_those_benchmarks/"> <img alt="Qwen3.8-Flash-Next: Time to Update Those Benchmarks" src="https://preview.redd.it/6sdkwxr3swlh1.png?width=640&crop=smart&auto=webp&a…
<!-- SC_OFF --><div class="md"><p>I'm running into all sorts of errors, has anybody got Qwen3.8-Flash-Next to work on a cluster of 2x DGX Sparks?</p> <p>If so could you post your settings/recipe?</p> <p>Thanks!</p> </div><!-- SC_ON -->   submitted by   <a href="https://ww…
<!-- SC_OFF --><div class="md"><p>Spent yesterday getting Qwen3.8 Flash and GLM 5.3 Flash up and running on my cluster of 4 x DGX Sparks with a view to replacing DeepSeek 0731... but.. really not that impressed with GLM 5.3 - overly verbose and takes for ever (was getting around …
<!-- SC_OFF --><div class="md"><p>How much system RAM? How much VRAM? How much SSD space?</p> <p>Ideally list for q3/4 but q2 might also work since I have seen 3.8 27B perform well even on q2.</p> <p>Currently I have 5070 Ti with 16GB VRAM and 48GB system RAM. I can upgrade syste…
Qwen3.8 Flash Next auf DGX Spark: MTP2 im Test MTP2-spekulative Dekodierung bringt Qwen3.8-Flash-Next auf der DGX Spark 15–24 % mehr Decode-Durchsatz bei 9 GB KV-Cache; Werkzeugqualität bleibt praktisch identisch. Der +43 %-Peak bei TG512 streut zu stark für ein belastbares Urtei…
RT @danieltvela: Ich habe Qwen3.8-Flash-Next auf einem PRO 6000 mit ~56 GB DRAM am Laufen. mehr auf Arint.info # AI # MTP # Performance # Prefilling # Qwen3 # vllm # arint_info https://x.com/danieltvela/status/2093999978473542089