PulseAugur
EN
LIVE 06:32:20

Qwen3.8-27B model shows strong performance across multiple hardware setups · 4 sources tracked

Users are reporting impressive performance and capabilities with the Qwen3.8-27B model across various hardware configurations. One user achieved a 262K context window on a single RTX 5090 using vLLM, demonstrating functional long-context handling with reasonable token generation speeds. Another setup utilizing a Strix Halo with an RTX 3090 Ti and llama.cpp achieved high token generation rates at 32K and 200K contexts, and notably outperformed a dual-RTX 3090 vLLM setup on the HumanEval benchmark. Further optimizations with DFlash2 and specific quantization levels on the Strix Halo showed that larger quantizations like Q5 can outperform Q4 for faster decoding due to better draft acceptance rates. AI

IMPACT Demonstrates advanced long-context capabilities and efficient inference techniques for local LLM deployments.

RANK_REASON User-generated reports on model performance and optimization techniques for local LLM inference.

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 4 sources. How we write summaries →

Qwen3.8-27B model shows strong performance across multiple hardware setups · 4 sources tracked

COVERAGE [4]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Fz1zz ·

    Single RTX 5090: Qwen3.8-27B NVFP4 at a real 262K context in vLLM — 77 tok/s short-context, 64.7 tok/s at 128K

    <!-- SC_OFF --><div class="md"><p>This is the Qwen3.8-27B setup I actually use every day on one RTX 5090. </p> <p>I wanted to write it down with enough detail that another 5090 owner can reproduce it instead of guessing which memory knobs I used.</p> <p>The short version: the ful…

  2. r/LocalLLaMA TIER_1 English(EN) · /u/TrifleHopeful5418 ·

    Qwen3.8-27B at 262K context on a Strix Halo + RTX 3090 Ti: 9.5 -> 153 tok/s, and it beats a dual-3090 vLLM box on HumanEval

    <!-- SC_OFF --><div class="md"><p>Spent a while treating layer placement, KV format and llama.cpp itself as experimental variables. 159 logged experiments. Numbers first, caveats after.</p> <p><strong>Hardware:</strong> AMD Ryzen AI MAX+ 395 (Strix Halo, 128 GB unified) + RTX 309…

  3. r/LocalLLaMA TIER_1 English(EN) · /u/stereohype ·

    Qwen3.8-27B (Q5_K_XL) on Strix Halo at 31 t/s decode: DFlash2 + Vulkan, the optimal setup

    <!-- SC_OFF --><div class="md"><p>Dense 27B, meet DFlash2. On my Flow Z13 (Ryzen AI Max+ 395, Radeon 8060S, 128GB), Qwen3.8-27B now decodes at <strong>31.4 t/s at 80W</strong> and prefills a 3k prompt at ~300 t/s. This is the dense followup to my DeepSeek V4 Flash guide from last…

  4. r/LocalLLaMA TIER_1 English(EN) · /u/seti_at_home ·

    Qwen3.8-27B Q8_0 on Strix Halo is seriously impressive

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vqme4y/qwen3827b_q8_0_on_strix_halo_is_seriously/"> <img alt="Qwen3.8-27B Q8_0 on Strix Halo is seriously impressive" src="https://external-preview.redd.it/MHI3NXJsOXVhd2poMXArTNwYs67Fb4dRjJDBvsZQ1H7SH3rcYPPn…