Alibaba's new open-source model, Qwen 3.8-27B, has outperformed Anthropic's Claude Opus 4.6 Max on the SWE-Bench Pro benchmark, achieving a score of 61.7 compared to 53.4. This 27-billion parameter model is capable of running on a single 24GB GPU and supports a native context window of 262,000 tokens. Qwen 3.8-27B is released under the Apache 2.0 license. AI
IMPACT This performance indicates strong competition in the open-source LLM space, potentially driving further innovation and accessibility.
RANK_REASON New benchmark result for an open-source model released by a major tech company. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Mastodon — mastodon.social →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →