An open-weight model, Qwen3.8 27B, has demonstrated superior performance compared to the proprietary Claude Opus 5. Qwen3.8 27B achieved a score of 52, while Claude Opus 5 scored 63.1, indicating an 11.1-point gap. Additionally, Qwen3.8 27B is noted to be eight times more cost-effective per million output tokens. AI
IMPACT Highlights the growing competitiveness of open-weight models against proprietary systems in terms of performance and cost-efficiency.
RANK_REASON Comparison of benchmark scores between an open-weight and a proprietary LLM. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Mastodon — mastodon.social →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →