A user on Reddit's r/LocalLLaMA subreddit is seeking feedback on the performance of the Qwen3.8-Flash-Next model running on their R9700 system. They are experiencing 863 tokens/s during prefill and 35 tokens/s during decoding with a context length of 230,000 tokens, and are unsure if these speeds are optimal for their hardware configuration. The user has detailed their setup, including the specific model version, backend, CPU, RAM, expert offloading, context settings, and the use of an ngram table, and is asking for advice on potential tuning to improve quality and speed. AI
IMPACT Provides insights into the practical performance limitations and tuning possibilities for local LLM deployments on consumer hardware.
RANK_REASON User-level inquiry about optimizing local LLM performance on specific hardware.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →