PulseAugur
EN
LIVE 22:12:03

Qwen3.8-Flash-Next performance questioned on R9700 system

A user on Reddit's r/LocalLLaMA subreddit is seeking feedback on the performance of the Qwen3.8-Flash-Next model running on their R9700 system. They are experiencing 863 tokens/s during prefill and 35 tokens/s during decoding with a context length of 230,000 tokens, and are unsure if these speeds are optimal for their hardware configuration. The user has detailed their setup, including the specific model version, backend, CPU, RAM, expert offloading, context settings, and the use of an ngram table, and is asking for advice on potential tuning to improve quality and speed. AI

IMPACT Provides insights into the practical performance limitations and tuning possibilities for local LLM deployments on consumer hardware.

RANK_REASON User-level inquiry about optimizing local LLM performance on specific hardware.

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Qwen3.8-Flash-Next performance questioned on R9700 system

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
User-level inquiry about optimizing local LLM performance on specific hardware.
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
8 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

Full methodology in our editorial standards.

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Designer_Elephant227 ·

    Qwen3.8-Flash-Next (5.05bpw + ngram at bf16) exl3 on one r9700: 863 t/s prefill and 35 t/s decode at 230k context (256k max), is that ok or am i missing something?

    <!-- SC_OFF --><div class="md"><p>Hi, i got QFN running on my single r9700 but im not sure if i did everything right to get the best quality and speed out of this setup. Dont want to annoy anybody, maybe someone with the same card can tell me if this looks normal.</p> <p>What i r…