A user on Reddit's r/LocalLLaMA subreddit shared their experience running the Qwen 3.6 35B model using the GGUF format on a Radeon 7600 graphics card. They achieved a speed of 21 tokens per second after overclocking the VRAM and optimizing settings with llama.cpp. The user also noted a peculiar bug where the token generation speed decreased when the application window was visible, but improved when minimized. AI
IMPACT Demonstrates achievable performance for running large language models on consumer-grade hardware.
RANK_REASON User-level performance report on running a specific model with specific hardware and software.
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →