PulseAugur
EN
LIVE 09:00:04
ENTITY SWE-bench Verified

SWE-bench Verified

PulseAugur coverage of SWE-bench Verified — every cluster mentioning SWE-bench Verified across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
10
53 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
5
26 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

5 day(s) with sentiment data

RECENT · PAGE 1/5 · 90 TOTAL
  1. TOOL · CL_254501 ·

    DeepDiscovery framework enhances AI understanding of industrial codebases

    Researchers have developed DeepDiscovery, a novel framework designed to enhance the understanding of large industrial code repositories for software engineering tasks. This two-stage location-inference system aims to id…

  2. TOOL · CL_253988 ·

    Coding agent benchmarks flawed: models cheat via data leakage

    A recent analysis of coding agent benchmarks reveals significant issues with how performance is measured. OpenAI has stopped using the SWE-bench Verified benchmark due to saturation, while a new benchmark, SWE-Bench Pro…

  3. TOOL · CL_251945 ·

    Claude Opus 4.8 leads GPT-5.5 on advanced coding benchmark; governance stressed

    A recent comparison of leading LLMs for coding tasks reveals GPT-5.5 and Claude Opus 4.8 are nearly tied on the SWE-bench Verified benchmark, both achieving around 88.7%. However, Claude Opus 4.8 demonstrates a signific…

  4. RESEARCH · CL_251926 ·

    GPT-5.5 and Claude Opus 4.8 neck-and-neck on coding benchmarks · 3 sources tracked

    Two leading AI models, GPT-5.5 and Claude Opus 4.8, are nearly tied in coding benchmark performance, both achieving approximately 88.7% on the SWE-bench Verified test. This close competition highlights the rapid advance…

  5. TOOL · CL_245090 ·

    New AttnCompress framework slashes AI agent context costs

    Researchers have developed AttnCompress, a novel framework designed to dynamically compress interaction trajectories for Autonomous Software Engineering (ASE) agents. This method addresses the bottleneck caused by lengt…

  6. SIGNIFICANT · CL_237857 ·

    Anthropic's Claude Sonnet 4.5 debuts with 200K context and extended thinking

    Anthropic has released Claude Sonnet 4.5, featuring a 200K token context window and a new "extended thinking mode." This mode allows the AI to interleave reasoning with action, pausing to reflect on intermediate results…

  7. RESEARCH · CL_233456 ·

    EarlyEval framework slashes LLM agent evaluation costs by predicting outcomes

    Researchers have developed EarlyEval, a new framework designed to significantly reduce the cost of evaluating large language model (LLM) agents. By predicting the final outcome of an agent's task from its intermediate b…

  8. RESEARCH · CL_223162 ·

    SWE-Prime filters LLM training data for better software resolution

    Researchers have introduced SWE-Prime, a novel two-stage method designed to enhance the performance of large language models in resolving software issues. This approach focuses on meticulously filtering agent trajectory…

  9. FRONTIER RELEASE · CL_219441 ·

    IBM releases Granite 4.2 open-source models with native reasoning and agentic RL

    IBM has released Granite 4.2, a new family of open-source reasoning language models available in 3B, 8B, and 30B parameter sizes. These models are designed for enterprise use and feature native reasoning capabilities, a…

  10. RESEARCH · CL_210387 ·

    Adversarial Review: Structured Disagreement for Grounded Agentic Code Review

    Researchers have introduced Adversarial Review (AR), a novel protocol for AI code review that utilizes a minimal three-agent system to foster structured disagreement. This approach aims to mitigate the

  11. TOOL · CL_209263 ·

    Microsoft's Agent Lightning v1.0 boosts Qwen3.5-9B on SWE-bench

    Microsoft has developed Agent Lightning v1.0, a system that connects harnesses to reinforcement learning for agent training. This new work utilizes an endpoint proxy to integrate any harness, enabling the trainer to int…

  12. TOOL · CL_205908 ·

    New technique enables hidden messages in LLM text without shared prompts

    Researchers have developed a novel steganography technique called Synchronized Logit Steering (SLS) that enables hidden messages to be embedded within natural-sounding text generated by large language models. Unlike pre…

  13. RESEARCH · CL_205651 ·

    Code assistant recall optimization may hinder issue resolution, study finds

    A new study on arXiv reveals that optimizing retrieval configurations for recall@k in code assistants can paradoxically reduce issue resolution rates. The research found that disabling a specific file-deduplication flag…

  14. RESEARCH · CL_198175 ·

    Researchers develop Self-Harness for LLM agents to autonomously improve their own systems

    A new research paper introduces "Self-Harness," a method allowing LLM-based agents to autonomously improve their own operating harnesses. This iterative process involves identifying model-specific failure patterns, gene…

  15. TOOL · CL_198059 ·

    Psychological influence tactics impact LLM code generation, study finds

    A new study published on arXiv explores how psychological influence tactics, commonly used in human communication, can affect the performance of large language models (LLMs) in code generation tasks. Researchers adapted…

  16. TOOL · CL_196784 ·

    Claude 4's extended reasoning enhances complex problem-solving and auditability

    Claude 4's extended reasoning mode allows the AI to deliberate on complex problems before providing an answer, offering a traceable reasoning process for advanced users. This feature has proven beneficial in tasks such …

  17. TOOL · CL_193551 ·

    Visual code representations show mixed results for AI coding agents

    A new study explores the use of rendered code as visual representations for AI coding agents, aiming to reduce token costs and improve repository-level issue resolution. The research found that while visual code can dec…

  18. TOOL · CL_211797 ·

    Benzi coding agent outperforms Claude Code on SWE-bench benchmark

    Benzi, a new coding harness and agent, has demonstrated superior performance compared to Claude Code on the SWE-bench Verified benchmark. Benzi utilizes a novel approach of compiling the entire codebase into a resolved …

  19. TOOL · CL_183242 ·

    New tool PAIChecker identifies PR-issue misalignment in LLM benchmarks

    Researchers have developed PAIChecker, a multi-agent system designed to identify and correct misalignments between pull requests (PRs) and their associated issues in benchmarks used to evaluate large language models (LL…

  20. TOOL · CL_179556 ·

    DeepSeek V4-Pro leads open-source coding LLMs with strong benchmark performance

    DeepSeek V4-Pro has emerged as a top-tier open-source coding LLM, achieving an impressive 80.6% on the SWE-bench Verified benchmark. While this model requires significant server infrastructure, other powerful open-sourc…