Cognition's new SWE-2 model excels on easy benchmarks, achieving near-perfect scores and offering a strong price-performance ratio for simpler tasks. However, a deeper analysis of its performance on harder, adversarial tasks reveals a significant drop-off, with a 65.5-point collapse on the Terminal-Bench TB4 compared to its TB2.1 score. This stark contrast suggests SWE-2's strengths lie in the "easy slice" of workloads, while its performance on genuinely challenging problems lags behind competitors like Fable 5.1 and GPT-6 Astra, indicating a potential training artifact that prioritizes cost-efficiency over robust performance on difficult edge cases. AI
IMPACT Highlights the importance of analyzing model performance across different difficulty levels, suggesting that easy-slice benchmarks may not reflect true capability on complex tasks.
RANK_REASON The item analyzes the performance of a released model (SWE-2) but focuses on interpretation and critique rather than a direct announcement or benchmark result.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →