Cognition 发布了其新的编码模型 SWE-2,该模型采用专家混合(MoE)架构,拥有 2.8 万亿参数,每个 token 激活 1040 亿参数。据报道,该模型在 Terminal-Bench 2.1 基准测试中取得了 92.8 分,在 FrontierCode 等任务上表现强劲,但在处理更长周期的代理任务方面落后于 Claude Fable 5.1 和 GPT-6 Astra 等竞争对手。SWE-2 已集成到 Cognition 的 Devin 编码助手产品中,但不能作为独立 API 或本地部署。 AI
影响 在 Terminal-Bench 2.1 上设定了新的 SOTA,但突显了长周期代理任务中仍然存在的差距。
排序理由 Frontier-lab 模型发布,附带系统卡和基准数据。
在 Mastodon — mastodon.social 阅读 →
- cognition
- Terminal-Bench 2.1
- Claude Fable 5.1
- Devin CLI
- Devin Desktop
- GPT-6 Astra
- Mixture of Experts
- OpenRouter
- FrontierCode
- reinforcement learning
AI 生成摘要 · Google Gemini · 来自 7 个来源。 我们如何撰写摘要 →