A user on the r/LocalLLaMA subreddit is experiencing significant performance issues with the Qwen 3.8 27B model when running it through vLLM. Despite trying various vLLM versions, quantization methods (FP8, NVFP4), and configuration flags, the model exhibits excessively long reasoning times, taking minutes to respond to basic prompts. This contrasts with other models like Qwen 3.6 and DeepSeek V4 Flash, which perform much faster on the same hardware. The user is seeking insights into potential solutions, such as specific chat template tweaks or generation parameter adjustments, to resolve this "endless thinking" behavior. AI
IMPACT Potential performance bottleneck for Qwen 3.8 27B users on vLLM, impacting usability for local deployments.
RANK_REASON User-reported issue with a specific model and inference engine combination.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →