PulseAugur
EN
LIVE 02:15:08

Qwen 3.8 27B model shows mixed results on SlopCodeBench

A user on Reddit's r/LocalLLaMA shared benchmark results for the Qwen 3.8 27B model on the SlopCodeBench. The Qwen model performed poorly on strict checkpoints, indicating potential issues with codebase management when used autonomously. However, it showed better performance on core checkpoints, suggesting it could function effectively as a coding assistant with clear direction. AI

IMPACT Indicates potential limitations of Qwen 3.8 27B in autonomous code management, while highlighting its utility as a directed coding assistant.

RANK_REASON User-generated benchmark results for a specific model on a coding benchmark. [lever_c_demoted from research: ic=1 ai=1.0]

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Qwen 3.8 27B model shows mixed results on SlopCodeBench

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/corruptbytes ·

    Qwen 3.8 27B SlopCodeBench results

    <!-- SC_OFF --><div class="md"><p>Howdy, I'm back again - running my favorite benchmark (it's still unsaturated for the time being so might as well!)</p> <p>previous runs <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vbtiy7/deepseek_v4_flash_on_slopcodebench/">a</a> <a h…