PulseAugur
实时 04:00:47
English(EN) Single RTX 5090: Qwen3.8-27B NVFP4 at a real 262K context in vLLM — 77 tok/s short-context, 64.7 tok/s at 128K

Qwen3.8-27B模型在多种硬件配置下表现强劲 · 已追踪4个来源

用户报告称Qwen3.8-27B模型在各种硬件配置下展现出令人印象深刻的性能和能力。一位用户在使用vLLM的单块RTX 5090上实现了262K的上下文窗口,以合理的token生成速度展示了长上下文处理能力。另一套使用Strix Halo搭配RTX 3090 Ti和llama.cpp的配置,在32K和200K上下文下实现了高token生成速率,并且在HumanEval基准测试中显著优于双RTX 3090 vLLM配置。在Strix Halo上使用DFlash2和特定量化级别进行的进一步优化表明,像Q5这样更大的量化级别由于更好的草稿接受率,可以优于Q4实现更快的解码。 AI

影响 展示了本地LLM部署的高级长上下文能力和高效推理技术。

排序理由 用户生成的关于模型性能和本地LLM推理优化技术的报告。

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 4 个来源。 我们如何撰写摘要 →

Qwen3.8-27B模型在多种硬件配置下表现强劲 · 已追踪4个来源

报道来源 [4]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Fz1zz ·

    单张RTX 5090:vLLM中Qwen3.8-27B NVFP4实现真实262K上下文 — 短上下文77 tok/s,128K时64.7 tok/s

    <!-- SC_OFF --><div class="md"><p>This is the Qwen3.8-27B setup I actually use every day on one RTX 5090. </p> <p>I wanted to write it down with enough detail that another 5090 owner can reproduce it instead of guessing which memory knobs I used.</p> <p>The short version: the ful…

  2. r/LocalLLaMA TIER_1 English(EN) · /u/TrifleHopeful5418 ·

    Qwen3.8-27B 在 Strix Halo + RTX 3090 Ti 上实现 262K 上下文:9.5 -> 153 tok/s,并在 HumanEval 上超越双 3090 vLLM 盒子

    <!-- SC_OFF --><div class="md"><p>Spent a while treating layer placement, KV format and llama.cpp itself as experimental variables. 159 logged experiments. Numbers first, caveats after.</p> <p><strong>Hardware:</strong> AMD Ryzen AI MAX+ 395 (Strix Halo, 128 GB unified) + RTX 309…

  3. r/LocalLLaMA TIER_1 English(EN) · /u/stereohype ·

    Strix Halo 上的 Qwen3.8-27B (Q5_K_XL) 解码速度达 31 t/s:DFlash2 + Vulkan,最优配置

    <!-- SC_OFF --><div class="md"><p>Dense 27B, meet DFlash2. On my Flow Z13 (Ryzen AI Max+ 395, Radeon 8060S, 128GB), Qwen3.8-27B now decodes at <strong>31.4 t/s at 80W</strong> and prefills a 3k prompt at ~300 t/s. This is the dense followup to my DeepSeek V4 Flash guide from last…

  4. r/LocalLLaMA TIER_1 English(EN) · /u/seti_at_home ·

    Qwen3.8-27B Q8_0 在 Strix Halo 上表现令人印象深刻

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vqme4y/qwen3827b_q8_0_on_strix_halo_is_seriously/"> <img alt="Qwen3.8-27B Q8_0 on Strix Halo is seriously impressive" src="https://external-preview.redd.it/MHI3NXJsOXVhd2poMXArTNwYs67Fb4dRjJDBvsZQ1H7SH3rcYPPn…