Cognition has released its new coding model, SWE-2, which boasts a massive 2.8 trillion parameters with 104 billion active per token using a Mixture of Experts (MoE) architecture. The model reportedly achieves a 92.8 score on the Terminal-Bench 2.1 benchmark, and while it shows strong performance on tasks like FrontierCode, it lags behind competitors like Claude Fable 5.1 and GPT-6 Astra on longer-horizon agentic tasks. SWE-2 is integrated into Cognition's Devin coding assistant products but is not available as a standalone API or for local deployment. AI
IMPACT Sets new SOTA on Terminal-Bench 2.1, but highlights remaining gaps in long-horizon agentic tasks.
RANK_REASON Frontier-lab model release with system card and benchmark data.
Read on Mastodon — mastodon.social →
- cognition
- Terminal-Bench 2.1
- Claude Fable 5.1
- Devin CLI
- Devin Desktop
- GPT-6 Astra
- Mixture of Experts
- OpenRouter
- FrontierCode
- reinforcement learning
AI-generated summary · Google Gemini · from 7 sources. How we write summaries →