PulseAugur
EN
LIVE 10:04:24

Users seek hardware advice for faster Qwen3.6 35B model inference

A user on Reddit is seeking hardware configurations to achieve high inference speeds with the Qwen3.6 35B model. They are currently experiencing around 270-300 tokens/second for prefill and 30 tokens/second for decode on their AMD RX6600XT and Ryzen 7 5700X setup. The user notes that existing online benchmarks are inaccurate for their hardware and is looking for advice from others who have achieved faster performance, specifically aiming for 1000+ prefill and 100+ decode tokens/second. AI

IMPACT Provides insights into real-world hardware performance for running large language models locally.

RANK_REASON User discussion about hardware performance for a specific model, not a new release or significant industry event.

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Users seek hardware advice for faster Qwen3.6 35B model inference

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Mrinohk ·

    Those in the 1000+ prefill and 100+ decode range on Qwen3.6 35B at Q4, what hardware are you running?

    <!-- SC_OFF --><div class="md"><p>Trying to see what I can scrounge together bare minimum hardware requirements to get up to that rough speed. </p> <p>Right now I'm running an RX6600XT and Ryzen 7 5700X with 32GB of DDR4 at 3600MHZ. CachyOS, vanilla llama.cpp built with ROCm and …