A user benchmarked the Qwen3.8-27B model on an Apple M5 Max chip, comparing its performance with "thinking on xhigh" versus "thinking off." The "thinking on xhigh" setting resulted in 5.5 times more tokens processed and a 6x longer runtime. Disabling the "thinking" feature significantly impacted output quality, placing the model behind other models like Qwen3.6-35B-A3B in terms of tokens per second. AI
IMPACT Provides insights into the performance trade-offs of large language models on consumer-grade hardware.
RANK_REASON User benchmark of a specific model version on consumer hardware. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →