PulseAugur
实时 07:29:26
English(EN) How many tokens/second output are you getting with Qwen3.8-27B?

Reddit 讨论 Qwen3.8-27B 模型性能

在 r/LocalLLaMA 子版块的用户正在讨论 Qwen3.8-27B 模型的性能,特别是其 token 输出速度。一位用户报告称,在使用 Qwen3.8-27B-heretic-ara 模型,通过 llama.ccp 进行 Q5_K_M GGUF 量化后,在配备 3090 GPU、64 GB RAM 和 AMD 7950x CPU 的系统上,大约能达到每秒 30-32 个 token。 AI

影响 为本地运行开源模型的用户提供了关于实际性能的见解。

排序理由 Reddit 上关于模型性能的用户讨论。

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Reddit 讨论 Qwen3.8-27B 模型性能

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/CooLittleFonzies ·

    Qwen3.8-27B 的输出速度是多少 tokens/秒?

    <!-- SC_OFF --><div class="md"><p>Trying to get a feel for where I stand. If you can list your relevant hardware and model used, that would be awesome.</p> <p>Here's mine:</p> <p><strong>Model:</strong> Qwen3.8-27B-heretic-ara, Q5_K_M GGUF</p> <p><strong>T/s</strong>: ~30-32 t/se…