A user on Reddit's r/LocalLLaMA shared benchmark results for the Qwen 3.8 27B model on the SlopCodeBench. The Qwen model performed poorly on strict checkpoints, indicating potential issues with codebase management when used autonomously. However, it showed better performance on core checkpoints, suggesting it could function effectively as a coding assistant with clear direction. AI
IMPACT Indicates potential limitations of Qwen 3.8 27B in autonomous code management, while highlighting its utility as a directed coding assistant.
RANK_REASON User-generated benchmark results for a specific model on a coding benchmark. [lever_c_demoted from research: ic=1 ai=1.0]
- Claude Code
- DeepSeek V4 Flash
- GPT-5.6 Sol
- HumanLayer Fable, Sol, and Kimi Benchmark Subset
- HumanLayer Opus 5 Benchmark Subset
- Kimi K3
- Qwen
- SlopCodeBench
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →