PulseAugur
EN
LIVE 02:52:48

Qwen 3.6 35B model runs at 21 tokens/sec on Radeon 7600 GPU

A user on Reddit's r/LocalLLaMA subreddit shared their experience running the Qwen 3.6 35B model using the GGUF format on a Radeon 7600 graphics card. They achieved a speed of 21 tokens per second after overclocking the VRAM and optimizing settings with llama.cpp. The user also noted a peculiar bug where the token generation speed decreased when the application window was visible, but improved when minimized. AI

IMPACT Demonstrates achievable performance for running large language models on consumer-grade hardware.

RANK_REASON User-level performance report on running a specific model with specific hardware and software.

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 3 sources. How we write summaries →

Qwen 3.6 35B model runs at 21 tokens/sec on Radeon 7600 GPU

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
User-level performance report on running a specific model with specific hardware and software.
Source corroboration
3 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
47 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

Full methodology in our editorial standards.

COVERAGE [3]

  1. Towards AI TIER_1 English(EN) · Gian Luca Bailo, Ph.D. ·

    Qwen3.8–27B on Two Mid-Range GPUs, Measured on Release Day

    <h4><em>Everyone will publish the benchmark scores. Here is a less glamorous and more useful question: what does it actually cost to run on a home machine — and why that number is not up for negotiation.</em></h4><figure><img alt="Illustration of an open-frame desktop PC on a des…

  2. r/LocalLLaMA TIER_1 English(EN) · /u/Sweaty_Perception655 ·

    Running Qwen 3.6 35B A3B-Q8_0 gguf on a cheap radeon 7600 at 18 token/s * update increased to 21 t/s

    <!-- SC_OFF --><div class="md"><blockquote> <p><a href="https://www.reddit.com/r/LocalLLaMA/?f=flair_name%3A%22Discussion%22"></a>I also have 64 gb ddr4 ryzen 5600 Using llama.cpp Ubuntu distro</p> <p>Settings are as follows</p> <p>--n-gpu-layers 999 \</p> <p>--n-cpu-moe 36 \</p>…

  3. r/LocalLLaMA TIER_1 English(EN) · /u/Sweaty_Perception655 ·

    Running Qwen 3.5 35B A3B-Q8_0 gguf on a cheap radeon 7600 at 18 token/s

    <!-- SC_OFF --><div class="md"><p>I also have 64 gb ddr4 ryzen 5600 Using llama.cpp Ubuntu distro</p> <p>Settings are as follows</p> <p>--n-gpu-layers 999 \</p> <p>--n-cpu-moe 37 \</p> <p>--no-mmap \</p> <p>-ctk q8_0 \</p> <p>-ctv q8_0 \</p> <p>-fa 1 \</p> <p>-c 9000 \</p> </div>…