PulseAugur
实时 22:51:09
English(EN) 880 tok/s on one 5090 Qwen3.8-27B in 4-bit NVFP4, full 262k context

Qwen3.8-27B 模型使用 NInfer 引擎在单张 RTX 5090 上达到 880 tok/s

一位用户在使用单张 RTX 5090 GPU 运行 Qwen3.8-27B 模型时取得了令人印象深刻的性能,在 4 位 NVFP4 量化和完整的 262k 上下文下达到了每秒 880 个 token。该速度是通过使用 NInfer 引擎实现的,用户报告称其比 llama.cpp 等其他引擎快得多。基准测试结果表明,这款 NVFP4 量化模型在 HumanEval+ 和 AIME 等任务上的表现与整数量化版本相当,同时速度却显著更快。 AI

影响 展示了在消费级硬件上本地 LLM 推理的显著加速,可能降低高级模型使用的门槛。

排序理由 用户报告的在消费级硬件上使用新型引擎运行特定模型的性能基准测试。

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Qwen3.8-27B 模型使用 NInfer 引擎在单张 RTX 5090 上达到 880 tok/s

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Ond7 ·

    880 tok/s on one 5090 Qwen3.8-27B in 4-bit NVFP4, full 262k context

    <!-- SC_OFF --><div class="md"><p>Numbers first, on a single RTX 5090, Running CachyOS with COSMIC, and the entire desktop costs about 150 MB of VRAM.</p> <p>880 tok/s aggregate at 6 parallel requests (peaked at 967 on one run). 200+ tok/s single stream with MTP speculative decod…