A user on Reddit is seeking hardware configurations to achieve high inference speeds with the Qwen3.6 35B model. They are currently experiencing around 270-300 tokens/second for prefill and 30 tokens/second for decode on their AMD RX6600XT and Ryzen 7 5700X setup. The user notes that existing online benchmarks are inaccurate for their hardware and is looking for advice from others who have achieved faster performance, specifically aiming for 1000+ prefill and 100+ decode tokens/second. AI
IMPACT Provides insights into real-world hardware performance for running large language models locally.
RANK_REASON User discussion about hardware performance for a specific model, not a new release or significant industry event.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →