PulseAugur
实时 11:16:57
English(EN) Those in the 1000+ prefill and 100+ decode range on Qwen3.6 35B at Q4, what hardware are you running?

用户寻求硬件建议以加速 Qwen3.6 35B 模型推理

一位Reddit用户正在寻求硬件配置建议,以期在使用Qwen3.6 35B模型时获得高推理速度。他们目前在AMD RX6600XT和Ryzen 7 5700X的配置下,预填充速度约为270-300 token/秒,解码速度约为30 token/秒。用户指出,现有的在线基准测试对其硬件而言并不准确,因此正在寻求已实现更快性能的其他用户的建议,目标是达到1000+ token/秒的预填充和100+ token/秒的解码速度。 AI

影响 提供了关于本地运行大型语言模型的实际硬件性能的见解。

排序理由 用户讨论特定模型的硬件性能,而非新发布或重要的行业事件。

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

用户寻求硬件建议以加速 Qwen3.6 35B 模型推理

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Mrinohk ·

    在 Qwen3.6 35B 的 Q4 上,预填充超过 1000、解码超过 100 的各位,你们在用什么硬件?

    <!-- SC_OFF --><div class="md"><p>Trying to see what I can scrounge together bare minimum hardware requirements to get up to that rough speed. </p> <p>Right now I'm running an RX6600XT and Ryzen 7 5700X with 32GB of DDR4 at 3600MHZ. CachyOS, vanilla llama.cpp built with ROCm and …