SWE Bench Pro
PulseAugur coverage of SWE Bench Pro — every cluster mentioning SWE Bench Pro across labs, papers, and developer communities, ranked by signal.
- instance of Claude (Opus 4.8) 90%
- instance of Opus 4.8 90%
- used by GLM-5.2 90%
- instance of Terminal-Bench 2.1 90%
- instance of Zhipu AI 90%
- instance of Claude Fable-5 90%
- instance of Terminal-Bench 90%
- instance of MAI-Thinking 1 90%
- instance of MAI-Code-1-Flash 90%
- used by Minimax 90%
- instance of cubic metre 90%
- used by MAI-Code-1-Flash 90%
12 day(s) with sentiment data
Anthropic's focus on 'abstention' in Opus 4.8 will drive adoption for critical coding tasks
Opus 4.8's improved ability to abstain from answering when uncertain, rather than providing incorrect information, is a critical feature for complex coding tasks. This trait, highlighted in recent evidence, could lead to increased adoption of Claude Opus for high-stakes software development where accuracy and reliability are paramount.
SWE-Bench Pro scores are rapidly increasing, with multiple models surpassing 50%
Recent evidence shows MiniMax's M3 model achieving 59% and Microsoft's MAI-Code-1-Flash achieving 51% on SWE-Bench Pro. This indicates a significant upward trend in AI coding benchmark performance, with several models now breaking the 50% barrier.
MiniMax M3 may become a leading open-source alternative for coding tasks
MiniMax's M3 model has demonstrated strong performance on SWE-Bench Pro (59%) and Terminal Bench 2 (66%), coupled with a 1M token context window. If its accessibility and performance remain competitive, it could emerge as a preferred open-source option for developers seeking advanced coding assistance, potentially challenging proprietary models.
-
Sonnet 5 achieves 63.2% on SWE-bench Pro; OpenAI's Terra claims unverified benchmark score
A new AI model, Sonnet 5, has achieved a score of 63.2% on the SWE-bench Pro benchmark. Separately, OpenAI's model, Terra, reportedly scored 84.3% on the Terminal-Bench, though this claim is vendor-stated, preview-only,…
-
CalibForge system synthesizes challenging AI agent training tasks
Researchers have developed CalibForge, a system designed to synthesize and refine terminal tasks for training AI agents. This system uses adversarial solver calibration, employing strategies like multi-solver disagreeme…
-
Tsinghua University releases VeriLoop Coder-E1 for verifiable code repair
Researchers from Tsinghua University have open-sourced VeriLoop Coder-E1, a model designed for verifiable recursive self-improvement in code repair. Built upon the Qwen3.6-27B architecture, VeriLoop Coder-E1 utilizes an…
-
xAI releases Grok 4.5 trained on real developer workflows · 1 source tracked
xAI has released Grok 4.5, a 1.5-trillion-parameter Mixture-of-Experts model trained on real developer interaction data from the Cursor IDE. This unique training approach, which includes multi-file diffs and debugger se…
-
MindForge pipeline trains small LLMs for full software engineering lifecycle
Researchers have developed MindForge, an automated pipeline designed to train smaller language models in comprehensive software engineering tasks. This system converts open-source command-line programs into source-free …
-
MiniMax M3, GLM-5.2, Kimi K3: Choosing Open-Weight Models for Agents
A comparison of three open-weight models—MiniMax M3, GLM-5.2, and Kimi K3—highlights that leaderboard scores alone are insufficient for self-hosting decisions. The article emphasizes factors like VRAM requirements, lice…
-
Anthropic's Claude Opus 5 tops leaderboards, but users debate value and guardrails
Anthropic has released Claude Opus 5, which has achieved top rankings on several AI leaderboards, including SWE-bench and FrontierBench. A key innovation is the introduction of an 'effort' parameter in the API, allowing…
-
Zhipu AI's GLM-5.2 leads open-weight models with 1M context window · 1 source tracked
GLM-5.2, a new open-weight model from China's Zhipu AI, has been recognized as the top-performing open model as of July 2026. It achieved a score of 51 in the Intelligence Index v4.1 by Artificial Analysis, placing it f…
-
OpenAI, Moonshot, Anthropic launch flagship models; benchmarks show varied strengths
In a rapid succession of releases, OpenAI, Moonshot AI, and Anthropic have launched their latest flagship models: GPT-5.6 Sol, Kimi K3, and Claude Opus 5, respectively. While all three models offer substantial context w…
-
Anthropic's Claude Opus 5 offers major performance gains at reduced cost · 8 sources tracked
Anthropic has released Claude Opus 5, a new frontier-class model that offers significant performance improvements at a reduced cost compared to its predecessor, Opus 4.8. While Anthropic's official benchmarks highlight …
-
Google expands Gemini with cheaper models and wider agent access
Google has expanded its Gemini AI offerings with the release of three new models, including Gemini 3.6 Flash, designed for coding and agentic tasks. The company also made its personal AI agent, Gemini Spark, available t…
-
Anthropic's Claude Opus 5 matches Fable 5 performance at half the price · 10 sources tracked
Anthropic has released Claude Opus 5, a new model that rivals the performance of Fable 5 at half the price. Early evaluations and user anecdotes suggest Opus 5 excels in coding, complex reasoning, and agentic tasks, oft…
-
Poolside AI releases Laguna S 2.1, a compact coding model with 1M context
Poolside AI has released Laguna S 2.1, an 118B parameter Mixture-of-Experts model with 8B activated parameters and a 1M token context window. The model was developed in under nine weeks and demonstrates strong performan…
-
RAG fails to fix hallucination; new mid-tier model targets agent costs
Retrieval-augmented generation (RAG) has been criticized for not solving AI hallucination, instead shifting the problem to the retrieval stage. A new mid-tier model, priced at $2/M, has achieved 63.2% on the SWE-bench P…
-
Developer finds Claude Code slows sprints by 40% due to hidden time sinks
A developer found that using Claude Code, even with advanced models like Opus 4.7, actually slowed down their development sprints by approximately 40%. This slowdown was attributed to the time spent reading and verifyin…
-
Khidi bridges gap between developers and cheaper open-weight AI models
A new service called Khidi aims to bridge the gap between developers and the cost-effectiveness of open-weight AI models. The founder explains that while open models have become technically comparable to flagship offeri…
-
Anthropic ships Claude Opus 4.7, OpenAI counters with GPT-5.5 · 1 source tracked
Anthropic has released Claude Opus 4.7, featuring a significant improvement in agentic coding tasks with a SWE-bench Pro score increase to 64.3% and enhanced high-resolution vision capabilities. OpenAI followed shortly …
-
Anthropic releases Claude Sonnet 5 with improved agentic coding performance
Anthropic has launched Claude Sonnet 5, a new mid-tier AI model that demonstrates improved agentic capabilities. This latest iteration surpasses its predecessor, Sonnet 4.6, across all benchmarks, achieving a 63.2% scor…
-
Anthropic releases Claude Sonnet 5 with enhanced agentic capabilities
Anthropic has released Claude Sonnet 5, an updated mid-tier model that significantly improves agentic capabilities and performance over its predecessor, Sonnet 4.6. This new model demonstrates enhanced abilities in plan…
-
GPT-5.6, Claude Fable 5, Gemini 3, GLM-5.2: Top AI models compared · 1 source tracked
A comparison of four leading AI models in July 2026—GPT-5.6 Sol, Claude Fable 5, Gemini 3, and GLM-5.2—reveals no single winner, with each excelling in different areas. GPT-5.6 Sol leads in speed for long agent sessions…