PulseAugur
中
实时 21:47:17
English(EN) Qwen3.8-Flash-Next (5.05bpw + ngram at bf16) exl3 on one r9700: 863 t/s prefill and 35 t/s decode at 230k context (256k max), is that ok or am i missing something?

Qwen3.8-Flash-Next 在 R9700 系统上的性能受到质疑

Reddit 的 r/LocalLLaMA 子版块上一位用户正在就 Qwen3.8-Flash-Next 模型在其 R9700 系统上的性能寻求反馈。他们在使用 230,000 个 token 的上下文长度时,预填充速度为 863 token/s,解码速度为 35 token/s,并且不确定这些速度是否对其硬件配置是最佳的。用户详细介绍了他们的设置,包括具体的模型版本、后端、CPU、RAM、专家卸载、上下文设置以及 ngram 表的使用,并寻求关于潜在调优以提高质量和速度的建议。 AI

影响 提供了关于在消费级硬件上本地 LLM 部署的实际性能限制和调优可能性的见解。

排序理由 用户级别关于在特定硬件上优化本地 LLM 性能的咨询。

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Qwen3.8-Flash-Next 在 R9700 系统上的性能受到质疑

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
用户级别关于在特定硬件上优化本地 LLM 性能的咨询。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
8 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Designer_Elephant227 ·

    Qwen3.8-Flash-Next (5.05bpw + ngram at bf16) exl3 在一台 r9700 上:863 t/s 预填充和 35 t/s 解码,上下文 230k (最大 256k),这正常吗,还是我遗漏了什么?

    <!-- SC_OFF --><div class="md"><p>Hi, i got QFN running on my single r9700 but im not sure if i did everything right to get the best quality and speed out of this setup. Dont want to annoy anybody, maybe someone with the same card can tell me if this looks normal.</p> <p>What i r…