SWE-bench Verified
PulseAugur coverage of SWE-bench Verified — every cluster mentioning SWE-bench Verified across labs, papers, and developer communities, ranked by signal.
- instance of SWE Bench Pro 90%
- instance of Apache Software License 2.0 90%
- instance of SWE-Bench Multilingual 90%
- instance of Terminal-Bench V2 90%
- instance of AppWorld 90%
- used by Apache Software License 2.0 70%
- used by Terminal Bench 2.0 70%
- used by DagsHub 70%
- instance of Terminal Bench 2.0 70%
- instance of Terminal-Bench 2.1 70%
- affiliated with SWE Bench Pro 70%
- used by LiveCodeBench 70%
5 day(s) with sentiment data
-
DeepDiscovery framework enhances AI understanding of industrial codebases
Researchers have developed DeepDiscovery, a novel framework designed to enhance the understanding of large industrial code repositories for software engineering tasks. This two-stage location-inference system aims to id…
-
Coding agent benchmarks flawed: models cheat via data leakage
A recent analysis of coding agent benchmarks reveals significant issues with how performance is measured. OpenAI has stopped using the SWE-bench Verified benchmark due to saturation, while a new benchmark, SWE-Bench Pro…
-
Claude Opus 4.8 leads GPT-5.5 on advanced coding benchmark; governance stressed
A recent comparison of leading LLMs for coding tasks reveals GPT-5.5 and Claude Opus 4.8 are nearly tied on the SWE-bench Verified benchmark, both achieving around 88.7%. However, Claude Opus 4.8 demonstrates a signific…
-
GPT-5.5 and Claude Opus 4.8 neck-and-neck on coding benchmarks · 3 sources tracked
Two leading AI models, GPT-5.5 and Claude Opus 4.8, are nearly tied in coding benchmark performance, both achieving approximately 88.7% on the SWE-bench Verified test. This close competition highlights the rapid advance…
-
New AttnCompress framework slashes AI agent context costs
Researchers have developed AttnCompress, a novel framework designed to dynamically compress interaction trajectories for Autonomous Software Engineering (ASE) agents. This method addresses the bottleneck caused by lengt…
-
Anthropic's Claude Sonnet 4.5 debuts with 200K context and extended thinking
Anthropic has released Claude Sonnet 4.5, featuring a 200K token context window and a new "extended thinking mode." This mode allows the AI to interleave reasoning with action, pausing to reflect on intermediate results…
-
EarlyEval framework slashes LLM agent evaluation costs by predicting outcomes
Researchers have developed EarlyEval, a new framework designed to significantly reduce the cost of evaluating large language model (LLM) agents. By predicting the final outcome of an agent's task from its intermediate b…
-
SWE-Prime filters LLM training data for better software resolution
Researchers have introduced SWE-Prime, a novel two-stage method designed to enhance the performance of large language models in resolving software issues. This approach focuses on meticulously filtering agent trajectory…
-
IBM releases Granite 4.2 open-source models with native reasoning and agentic RL
IBM has released Granite 4.2, a new family of open-source reasoning language models available in 3B, 8B, and 30B parameter sizes. These models are designed for enterprise use and feature native reasoning capabilities, a…
-
Adversarial Review: Structured Disagreement for Grounded Agentic Code Review
Researchers have introduced Adversarial Review (AR), a novel protocol for AI code review that utilizes a minimal three-agent system to foster structured disagreement. This approach aims to mitigate the
-
Microsoft's Agent Lightning v1.0 boosts Qwen3.5-9B on SWE-bench
Microsoft has developed Agent Lightning v1.0, a system that connects harnesses to reinforcement learning for agent training. This new work utilizes an endpoint proxy to integrate any harness, enabling the trainer to int…
-
New technique enables hidden messages in LLM text without shared prompts
Researchers have developed a novel steganography technique called Synchronized Logit Steering (SLS) that enables hidden messages to be embedded within natural-sounding text generated by large language models. Unlike pre…
-
Code assistant recall optimization may hinder issue resolution, study finds
A new study on arXiv reveals that optimizing retrieval configurations for recall@k in code assistants can paradoxically reduce issue resolution rates. The research found that disabling a specific file-deduplication flag…
-
Researchers develop Self-Harness for LLM agents to autonomously improve their own systems
A new research paper introduces "Self-Harness," a method allowing LLM-based agents to autonomously improve their own operating harnesses. This iterative process involves identifying model-specific failure patterns, gene…
-
Psychological influence tactics impact LLM code generation, study finds
A new study published on arXiv explores how psychological influence tactics, commonly used in human communication, can affect the performance of large language models (LLMs) in code generation tasks. Researchers adapted…
-
Claude 4's extended reasoning enhances complex problem-solving and auditability
Claude 4's extended reasoning mode allows the AI to deliberate on complex problems before providing an answer, offering a traceable reasoning process for advanced users. This feature has proven beneficial in tasks such …
-
Visual code representations show mixed results for AI coding agents
A new study explores the use of rendered code as visual representations for AI coding agents, aiming to reduce token costs and improve repository-level issue resolution. The research found that while visual code can dec…
-
Benzi coding agent outperforms Claude Code on SWE-bench benchmark
Benzi, a new coding harness and agent, has demonstrated superior performance compared to Claude Code on the SWE-bench Verified benchmark. Benzi utilizes a novel approach of compiling the entire codebase into a resolved …
-
New tool PAIChecker identifies PR-issue misalignment in LLM benchmarks
Researchers have developed PAIChecker, a multi-agent system designed to identify and correct misalignments between pull requests (PRs) and their associated issues in benchmarks used to evaluate large language models (LL…
-
DeepSeek V4-Pro leads open-source coding LLMs with strong benchmark performance
DeepSeek V4-Pro has emerged as a top-tier open-source coding LLM, achieving an impressive 80.6% on the SWE-bench Verified benchmark. While this model requires significant server infrastructure, other powerful open-sourc…