PulseAugur
实时 18:09:04
English(EN) Qwen3.8-27B NVFP4 with vision + 451K token KV-cache on one RTX 5090 (power limited to 400W) at 120 tokens/s average

Qwen3.8-27B 模型在单张 RTX 5090 上实现 451K KV-cache

一位用户成功地在单张 RTX 5090 GPU 上配置了具备视觉能力和 451K token KV-cache 的 Qwen3.8-27B 模型。该设置利用 vLLM 和特定的 NVFP4 量化,即使在功率限制为 400W 的情况下,平均速度也能达到每秒 120 个 token。此配置在编码任务上表现准确,并支持最多三个并行会话。 AI

影响 展示了在消费级硬件上高效本地部署大上下文模型的可能性。

排序理由 用户在消费级硬件上对现有模型进行的配置和基准测试。

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Qwen3.8-27B 模型在单张 RTX 5090 上实现 451K KV-cache

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/t4a8945 ·

    Qwen3.8-27B NVFP4 视觉+451K token KV-cache 在单张 RTX 5090 (功率限制400W) 上实现平均120 tokens/s

    <!-- SC_OFF --><div class="md"><p>Hello,</p> <p>So I've been trying lots of combinations in that never-ending landscape of options and settings.</p> <p>I wanted a proper quant of 3.8 27B running as fast as possible on my 5090 at 400W, with vision and with as much KV-cache as poss…