PulseAugur
EN
LIVE 10:49:26
ENTITY SWE Bench Pro

SWE Bench Pro

PulseAugur coverage of SWE Bench Pro — every cluster mentioning SWE Bench Pro across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
23
89 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
4
23 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
TIMELINE
  1. 2026-07-08 research_milestone OpenAI audited the SWE-Bench Pro coding benchmark and found it unreliable. source
  2. 2026-07-08 research_milestone OpenAI audited SWE-Bench Pro and found it unreliable for measuring frontier coding capability. source
SENTIMENT · 30D

12 day(s) with sentiment data

LAB BRAIN
hypothesis resolved confirmed conf 0.70

Anthropic's focus on 'abstention' in Opus 4.8 will drive adoption for critical coding tasks

Opus 4.8's improved ability to abstain from answering when uncertain, rather than providing incorrect information, is a critical feature for complex coding tasks. This trait, highlighted in recent evidence, could lead to increased adoption of Claude Opus for high-stakes software development where accuracy and reliability are paramount.

observation resolved confirmed conf 0.85

SWE-Bench Pro scores are rapidly increasing, with multiple models surpassing 50%

Recent evidence shows MiniMax's M3 model achieving 59% and Microsoft's MAI-Code-1-Flash achieving 51% on SWE-Bench Pro. This indicates a significant upward trend in AI coding benchmark performance, with several models now breaking the 50% barrier.

hypothesis resolved confirmed conf 0.65

MiniMax M3 may become a leading open-source alternative for coding tasks

MiniMax's M3 model has demonstrated strong performance on SWE-Bench Pro (59%) and Terminal Bench 2 (66%), coupled with a 1M token context window. If its accessibility and performance remain competitive, it could emerge as a preferred open-source option for developers seeking advanced coding assistance, potentially challenging proprietary models.

All hypotheses →

RECENT · PAGE 1/5 · 89 TOTAL
  1. TOOL · CL_186078 ·

    Sonnet 5 achieves 63.2% on SWE-bench Pro; OpenAI's Terra claims unverified benchmark score

    A new AI model, Sonnet 5, has achieved a score of 63.2% on the SWE-bench Pro benchmark. Separately, OpenAI's model, Terra, reportedly scored 84.3% on the Terminal-Bench, though this claim is vendor-stated, preview-only,…

  2. RESEARCH · CL_186963 ·

    CalibForge system synthesizes challenging AI agent training tasks

    Researchers have developed CalibForge, a system designed to synthesize and refine terminal tasks for training AI agents. This system uses adversarial solver calibration, employing strategies like multi-solver disagreeme…

  3. TOOL · CL_178119 ·

    Tsinghua University releases VeriLoop Coder-E1 for verifiable code repair

    Researchers from Tsinghua University have open-sourced VeriLoop Coder-E1, a model designed for verifiable recursive self-improvement in code repair. Built upon the Qwen3.6-27B architecture, VeriLoop Coder-E1 utilizes an…

  4. SIGNIFICANT · CL_175218 ·

    xAI releases Grok 4.5 trained on real developer workflows · 1 source tracked

    xAI has released Grok 4.5, a 1.5-trillion-parameter Mixture-of-Experts model trained on real developer interaction data from the Cursor IDE. This unique training approach, which includes multi-file diffs and debugger se…

  5. RESEARCH · CL_171855 ·

    MindForge pipeline trains small LLMs for full software engineering lifecycle

    Researchers have developed MindForge, an automated pipeline designed to train smaller language models in comprehensive software engineering tasks. This system converts open-source command-line programs into source-free …

  6. TOOL · CL_164320 ·

    MiniMax M3, GLM-5.2, Kimi K3: Choosing Open-Weight Models for Agents

    A comparison of three open-weight models—MiniMax M3, GLM-5.2, and Kimi K3—highlights that leaderboard scores alone are insufficient for self-hosting decisions. The article emphasizes factors like VRAM requirements, lice…

  7. SIGNIFICANT · CL_162820 ·

    Anthropic's Claude Opus 5 tops leaderboards, but users debate value and guardrails

    Anthropic has released Claude Opus 5, which has achieved top rankings on several AI leaderboards, including SWE-bench and FrontierBench. A key innovation is the introduction of an 'effort' parameter in the API, allowing…

  8. SIGNIFICANT · CL_162724 ·

    Zhipu AI's GLM-5.2 leads open-weight models with 1M context window · 1 source tracked

    GLM-5.2, a new open-weight model from China's Zhipu AI, has been recognized as the top-performing open model as of July 2026. It achieved a score of 51 in the Intelligence Index v4.1 by Artificial Analysis, placing it f…

  9. FRONTIER RELEASE · CL_162247 ·

    OpenAI, Moonshot, Anthropic launch flagship models; benchmarks show varied strengths

    In a rapid succession of releases, OpenAI, Moonshot AI, and Anthropic have launched their latest flagship models: GPT-5.6 Sol, Kimi K3, and Claude Opus 5, respectively. While all three models offer substantial context w…

  10. FRONTIER RELEASE · CL_162242 ·

    Anthropic's Claude Opus 5 offers major performance gains at reduced cost · 8 sources tracked

    Anthropic has released Claude Opus 5, a new frontier-class model that offers significant performance improvements at a reduced cost compared to its predecessor, Opus 4.8. While Anthropic's official benchmarks highlight …

  11. SIGNIFICANT · CL_161792 ·

    Google expands Gemini with cheaper models and wider agent access

    Google has expanded its Gemini AI offerings with the release of three new models, including Gemini 3.6 Flash, designed for coding and agentic tasks. The company also made its personal AI agent, Gemini Spark, available t…

  12. FRONTIER RELEASE · CL_160596 ·

    Anthropic's Claude Opus 5 matches Fable 5 performance at half the price · 10 sources tracked

    Anthropic has released Claude Opus 5, a new model that rivals the performance of Fable 5 at half the price. Early evaluations and user anecdotes suggest Opus 5 excels in coding, complex reasoning, and agentic tasks, oft…

  13. SIGNIFICANT · CL_155809 ·

    Poolside AI releases Laguna S 2.1, a compact coding model with 1M context

    Poolside AI has released Laguna S 2.1, an 118B parameter Mixture-of-Experts model with 8B activated parameters and a 1M token context window. The model was developed in under nine weeks and demonstrates strong performan…

  14. COMMENTARY · CL_153153 ·

    RAG fails to fix hallucination; new mid-tier model targets agent costs

    Retrieval-augmented generation (RAG) has been criticized for not solving AI hallucination, instead shifting the problem to the retrieval stage. A new mid-tier model, priced at $2/M, has achieved 63.2% on the SWE-bench P…

  15. COMMENTARY · CL_146530 ·

    Developer finds Claude Code slows sprints by 40% due to hidden time sinks

    A developer found that using Claude Code, even with advanced models like Opus 4.7, actually slowed down their development sprints by approximately 40%. This slowdown was attributed to the time spent reading and verifyin…

  16. TOOL · CL_145206 ·

    Khidi bridges gap between developers and cheaper open-weight AI models

    A new service called Khidi aims to bridge the gap between developers and the cost-effectiveness of open-weight AI models. The founder explains that while open models have become technically comparable to flagship offeri…

  17. SIGNIFICANT · CL_143170 ·

    Anthropic ships Claude Opus 4.7, OpenAI counters with GPT-5.5 · 1 source tracked

    Anthropic has released Claude Opus 4.7, featuring a significant improvement in agentic coding tasks with a SWE-bench Pro score increase to 64.3% and enhanced high-resolution vision capabilities. OpenAI followed shortly …

  18. SIGNIFICANT · CL_140905 ·

    Anthropic releases Claude Sonnet 5 with improved agentic coding performance

    Anthropic has launched Claude Sonnet 5, a new mid-tier AI model that demonstrates improved agentic capabilities. This latest iteration surpasses its predecessor, Sonnet 4.6, across all benchmarks, achieving a 63.2% scor…

  19. SIGNIFICANT · CL_140840 ·

    Anthropic releases Claude Sonnet 5 with enhanced agentic capabilities

    Anthropic has released Claude Sonnet 5, an updated mid-tier model that significantly improves agentic capabilities and performance over its predecessor, Sonnet 4.6. This new model demonstrates enhanced abilities in plan…

  20. COMMENTARY · CL_137312 ·

    GPT-5.6, Claude Fable 5, Gemini 3, GLM-5.2: Top AI models compared · 1 source tracked

    A comparison of four leading AI models in July 2026—GPT-5.6 Sol, Claude Fable 5, Gemini 3, and GLM-5.2—reveals no single winner, with each excelling in different areas. GPT-5.6 Sol leads in speed for long agent sessions…