Independent benchmarks reveal that Qwen3.5 2B, a non-reasoning model, achieves 43.8% on the GPQA benchmark. However, its performance significantly drops on more complex tasks, scoring only 5% on HLE, 15% on Long Context Reasoning, and 7.2% on SciCode. This indicates a substantial performance gap on challenging evaluations. AI
IMPACT Highlights performance limitations of smaller open-source models on complex reasoning tasks.
RANK_REASON Independent benchmark results for an open-source model. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Mastodon — mastodon.social →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →