PulseAugur
EN
LIVE 13:08:18
ENTITY SWE Bench Pro

SWE Bench Pro

PulseAugur coverage of SWE Bench Pro — every cluster mentioning SWE Bench Pro across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
8
57 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
5
17 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
TIMELINE
  1. 2026-07-08 research_milestone OpenAI audited the SWE-Bench Pro coding benchmark and found it unreliable. source
  2. 2026-07-08 research_milestone OpenAI audited SWE-Bench Pro and found it unreliable for measuring frontier coding capability. source
SENTIMENT · 30D

5 day(s) with sentiment data

LAB BRAIN
hypothesis resolved confirmed conf 0.70

Anthropic's focus on 'abstention' in Opus 4.8 will drive adoption for critical coding tasks

Opus 4.8's improved ability to abstain from answering when uncertain, rather than providing incorrect information, is a critical feature for complex coding tasks. This trait, highlighted in recent evidence, could lead to increased adoption of Claude Opus for high-stakes software development where accuracy and reliability are paramount.

observation resolved confirmed conf 0.85

SWE-Bench Pro scores are rapidly increasing, with multiple models surpassing 50%

Recent evidence shows MiniMax's M3 model achieving 59% and Microsoft's MAI-Code-1-Flash achieving 51% on SWE-Bench Pro. This indicates a significant upward trend in AI coding benchmark performance, with several models now breaking the 50% barrier.

hypothesis resolved confirmed conf 0.65

MiniMax M3 may become a leading open-source alternative for coding tasks

MiniMax's M3 model has demonstrated strong performance on SWE-Bench Pro (59%) and Terminal Bench 2 (66%), coupled with a 1M token context window. If its accessibility and performance remain competitive, it could emerge as a preferred open-source option for developers seeking advanced coding assistance, potentially challenging proprietary models.

All hypotheses →

RECENT · PAGE 1/6 · 104 TOTAL
  1. COMMENTARY · CL_257920 ·

    GPT-6 Astra and Claude Fable 5.1 face off in benchmark showdown

    A direct comparison between OpenAI's GPT-6 Astra and Anthropic's Claude Fable 5.1 reveals distinct strengths for each advanced AI model. GPT-6 Astra excels in handling massive codebases with its 1.1M context window and …

  2. RESEARCH · CL_256472 ·

    OpenAI flags SWE-bench Pro coding benchmark flaws; new research reframes NN training

    OpenAI has identified significant reliability issues within the SWE-bench Pro coding benchmark, a tool it had previously endorsed as a replacement for the contaminated SWE-bench benchmark. This new analysis suggests tha…

  3. TOOL · CL_251933 ·

    MetaRSI-v1 advances AI self-improvement capabilities · 1 source tracked

    CosmosMind, in collaboration with several universities, has introduced MetaRSI-v1, a novel meta-recursive architecture designed to improve the process of recursive self-improvement (RSI) in AI models. This new framework…

  4. TOOL · CL_251945 ·

    Claude Opus 4.8 leads GPT-5.5 on advanced coding benchmark; governance stressed

    A recent comparison of leading LLMs for coding tasks reveals GPT-5.5 and Claude Opus 4.8 are nearly tied on the SWE-bench Verified benchmark, both achieving around 88.7%. However, Claude Opus 4.8 demonstrates a signific…

  5. RESEARCH · CL_249076 ·

    ByteDance's HarnessDev benchmark tests LLMs' ability to build agent code

    Researchers from ByteDance Seed and other institutions have introduced HarnessDev, a new benchmark designed to evaluate an LLM's ability to create its own agent harnesses. Unlike traditional benchmarks that fix the harn…

  6. TOOL · CL_226646 ·

    New Agentic Coding Index ranks LLMs by intelligence density

    A Reddit user has developed a new metric called the Agentic Coding Index (ACI) to evaluate Large Language Models (LLMs) on coding tasks. The ACI aggregates scores from several coding benchmarks, including SWE-bench Pro,…

  7. RESEARCH · CL_223162 ·

    SWE-Prime filters LLM training data for better software resolution

    Researchers have introduced SWE-Prime, a novel two-stage method designed to enhance the performance of large language models in resolving software issues. This approach focuses on meticulously filtering agent trajectory…

  8. FRONTIER RELEASE · CL_219957 ·

    Alibaba previews Qwen4 architecture with cost-efficient Qwen3.8-Flash-Next model

    Alibaba's Qwen team has released Qwen3.8-Flash-Next, an open-weight multimodal MoE model that previews the architecture for the upcoming Qwen4. This new model boasts significant cost-efficiency, activating only 6B param…

  9. TOOL · CL_218759 ·

    Microsoft's AutoSaddler automates AI agent harness optimization

    Researchers from Microsoft have developed AutoSaddler, a novel system designed to automatically optimize AI agent harnesses. This method treats the harness as code and learns to patch it offline using failure traces. Au…

  10. COMMENTARY · CL_218458 ·

    Anthropic's Fable 5 struggles as businesses favor cheaper Opus 4.8

    Despite developing highly capable AI models like Fable 5, Anthropic is experiencing a shift in enterprise spending towards its older, more affordable Opus 4.8 model. Data from Ramp indicates that Opus 4.8 accounts for a…

  11. TOOL · CL_216420 ·

    LLM production readiness requires 10 tests beyond benchmarks

    An LLM evaluation checklist for 2026 emphasizes moving beyond simple benchmark scores to comprehensive production readiness testing. The approach, developed by Quokka Labs, focuses on evaluating the entire AI system, in…

  12. SIGNIFICANT · CL_210098 ·

    Qwen3.8-27B open-source model tops leaderboards with advanced agency

    Alibaba's Qwen team has open-sourced Qwen3.8-27B, a 27-billion parameter model that has achieved top rankings on several benchmarks, including SWE Bench Pro and OSWorld. This model demonstrates significant advancements …

  13. SIGNIFICANT · CL_207137 ·

    九章智算云 focuses on training-inference consistency for AI infrastructure

    九章智算云 is developing an AI infrastructure system focused on "training-inference consistency" to support the increasing reliance on reinforcement learning (RL) for scaling model capabilities. This system aims to efficient…

  14. TOOL · CL_202754 ·

    Alibaba's Qwen 3.8-27B model surpasses Claude Opus 4.6 Max on SWE-Bench Pro

    Alibaba's new open-source model, Qwen 3.8-27B, has outperformed Anthropic's Claude Opus 4.6 Max on the SWE-Bench Pro benchmark, achieving a score of 61.7 compared to 53.4. This 27-billion parameter model is capable of r…

  15. SIGNIFICANT · CL_201694 ·

    Qwen3.8-27B open-source model rivals Claude Opus on benchmarks

    Qwen has released its Qwen3.8-27B model, an open-source model that rivals the performance of larger, proprietary models like Claude Opus 4.6 Max on various benchmarks, particularly in coding and agent tasks. This 27-bil…

  16. TOOL · CL_186078 ·

    Sonnet 5 achieves 63.2% on SWE-bench Pro; OpenAI's Terra claims unverified benchmark score

    A new AI model, Sonnet 5, has achieved a score of 63.2% on the SWE-bench Pro benchmark. Separately, OpenAI's model, Terra, reportedly scored 84.3% on the Terminal-Bench, though this claim is vendor-stated, preview-only,…

  17. RESEARCH · CL_186963 ·

    CalibForge system synthesizes challenging AI agent training tasks

    Researchers have developed CalibForge, a system designed to synthesize and refine terminal tasks for training AI agents. This system uses adversarial solver calibration, employing strategies like multi-solver disagreeme…

  18. TOOL · CL_178119 ·

    Tsinghua University releases VeriLoop Coder-E1 for verifiable code repair

    Researchers from Tsinghua University have open-sourced VeriLoop Coder-E1, a model designed for verifiable recursive self-improvement in code repair. Built upon the Qwen3.6-27B architecture, VeriLoop Coder-E1 utilizes an…

  19. SIGNIFICANT · CL_175218 ·

    xAI releases Grok 4.5 trained on real developer workflows · 1 source tracked

    xAI has released Grok 4.5, a 1.5-trillion-parameter Mixture-of-Experts model trained on real developer interaction data from the Cursor IDE. This unique training approach, which includes multi-file diffs and debugger se…

  20. RESEARCH · CL_171855 ·

    MindForge pipeline trains small LLMs for full software engineering lifecycle

    Researchers have developed MindForge, an automated pipeline designed to train smaller language models in comprehensive software engineering tasks. This system converts open-source command-line programs into source-free …