A user tested the Qwen3.5-9B model on an Apple M1 Max with 64GB of RAM, using Ollama for local execution. While the model's prompt suggested it could outperform GPT-4 in Japanese, the test focused on actual performance metrics. The Qwen3.5-9B model, quantized to Q4 and using approximately 6.6GB of VRAM, produced a 131-character Japanese response but generated a total of 3936 tokens. The vast majority of these tokens were internal "thinking" tokens, not visible to the user, which significantly impacted the perceived speed. The user found that relying solely on the `eval rate` metric could be misleading, as the actual user experience was much slower than the token generation rate suggested. AI
IMPACT Highlights how internal "thinking" tokens can inflate output and mislead performance metrics for local LLMs, impacting user experience and expectations.
RANK_REASON User benchmark and performance analysis of a specific LLM. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →