SWE-bench
PulseAugur coverage of SWE-bench — every cluster mentioning SWE-bench across labs, papers, and developer communities, ranked by signal.
- used by GLM-5.2 90%
- instance of Claude Sonnet 4.6 90%
- instance of DeepSeek V4-Pro 90%
- used by arXiv 70%
- instance of HumanEval 70%
- instance of Terminal-Bench 70%
- competes with Claude Opus 4-8 70%
- instance of SWE-bench Verified 70%
- instance of Claude Fable-5 70%
- competes with Claude Fable-5 70%
- used by vLLM 70%
- used by LiveCodeBench 70%
7 day(s) with sentiment data
-
AI agents improve coding and robotics with new RL and framework
A new reinforcement learning method called CompactionRL has been developed by researchers at Tsinghua University, which improves the performance of AI coding agents on the SWE-bench benchmark by 7 points. This method en…
-
New method uses internal LLM representations to detect reward hacking
A new research paper introduces a method for detecting and understanding "reward hacking" in large language models by analyzing their internal representations. The study found that simple "difference of means" (DoM) vec…
-
OpenAI flags SWE-bench Pro coding benchmark flaws; new research reframes NN training
OpenAI has identified significant reliability issues within the SWE-bench Pro coding benchmark, a tool it had previously endorsed as a replacement for the contaminated SWE-bench benchmark. This new analysis suggests tha…
-
Paper questions LLM-agent leaderboard validity
A new paper questions the validity of LLM-agent leaderboards, arguing that direct comparisons of ranked agents can be misleading. The authors highlight that differences in task mixtures, data sources, release details, a…
-
AI agents explore recursive self-improvement with new frameworks · 7 sources tracked
Researchers are exploring recursive self-improvement (RSI) in AI agents, a process where AI systems enhance their own capabilities and the methods for future improvement. New frameworks like Dream-RSI and RSIAgent lever…
-
New benchmark evaluates cost-effective AI for code repository exploration
Researchers have developed a new benchmark, IssueLoc-Bench, to evaluate the cost-effectiveness of different AI models for repository exploration in coding agent pipelines. The study found that while higher-quality explo…
-
HARTS system accelerates agentic reinforcement learning for hybrid-attention models
Researchers have developed HARTS, a novel system designed to improve the efficiency of agentic reinforcement learning (RL) for hybrid-attention models. HARTS addresses the challenge of recomputing shared prefixes in irr…
-
HY4 language model achieves 1-bit quantization with minimal accuracy loss
A new 1-bit quantization for the HY4 language model has been released, showing promising results with minimal accuracy loss compared to BF16. The quantization, which was initially mislabeled as Q1 but is actually 2.38-b…
-
New research explores robust, adaptive, and structured reinforcement learning techniques · 10 sources tracked
Multiple research papers published on arXiv explore advancements in reinforcement learning (RL) techniques. One paper unifies regularization-based methods for robust deep RL against adversarial perturbations, proposing …
-
Fireworks AI touts DeepSeek V4 Pro's performance and cost advantages
Fireworks AI has announced that its platform, utilizing the DeepSeek V4 Pro model, outperforms Anthropic's Claude Fable 5 on SWE-Bench and LiveCodeBench benchmarks. The DeepSeek V4 Pro offers a lower cost per solved tas…
-
Agent Engine's Architecture Thwarts Prompt Injection Attempts
An open-source agent engine called PlannerCritic, designed with a two-LLM architecture for planning and review, successfully resisted prompt injection attempts. The engine's design, which includes deterministic gates th…
-
Muse Glimmer 30B: Community tests reveal mixed performance vs. benchmarks
Two weeks after Meta's release of the Muse Glimmer 30B model, community testing has revealed a more nuanced performance compared to official benchmarks. While Glimmer excels in agentic tasks and tool calling, independen…
-
Meta's Muse Video leaks, Replit offers free GPT-5.6 Luna, and AI regulation debated
Meta's Muse Video model has been revealed in a closed beta, showcasing its ability to generate 10-second videos with strong temporal consistency and detail, including native audio. Meanwhile, Replit has launched a Free …
-
Ornith AI releases Ornith-1.5 model family with self-improvement focus
Ornith AI has released the Ornith-1.5 model family, featuring a 9B dense model and 35B and 397B Mixture-of-Experts (MoE) variants. These models are designed for self-improvement and have demonstrated competitive perform…
-
Coding benchmark scores may not reflect general AI capability, study finds
A new paper argues that optimizing AI models for specific coding benchmarks like SWE-bench does not necessarily improve their general coding capabilities. Researchers found that models trained on these benchmarks showed…
-
HELIX system enables recursive self-improvement for AI agents
Researchers have introduced HELIX, a new substrate designed for the co-evolution of AI models and their runtime harnesses. This approach focuses on improving the harness—the system that mediates context, tools, and cont…
-
Graft tool enhances coding agents by building codebase graphs
Graft is a new tool designed to enhance the contextual understanding of coding agents like Claude Code, Cursor, and Gemini. It achieves this by building a graph of the codebase, represented as linked markdown files, whi…
-
DeepSeek V4 Pro shows strong coding gains in practical frontend tests
DeepSeek has released its V4 Pro model, which shows significant improvements in code generation and reasoning capabilities compared to its predecessor, V3.1. While official benchmarks highlight gains, this article focus…
-
Meta releases open-weight Muse Glimmer model for agentic tasks
Meta has released Muse Glimmer, a new 30B parameter open-weight model licensed under Apache 2.0. The model is designed for end-to-end agentic task completion, reliable tool use, and multi-step reasoning, showing strong …
-
Developer creates custom harness to benchmark coding LLMs on personal bugs
A developer has created a Python-based harness to evaluate coding LLMs against a personal corpus of bugs, rather than relying on public benchmarks like SWE-bench. This approach aims to provide more relevant performance …