SWE-rebench
PulseAugur coverage of SWE-rebench — every cluster mentioning SWE-rebench across labs, papers, and developer communities, ranked by signal.
- 2026-07-01 research_milestone The SWE-rebench leaderboard was updated with new models and an improved UI for comparing AI performance on coding tasks. source
1 day(s) with sentiment data
-
SWE-rebench adds multilingual coding tasks, GLM-5.2 leads leaderboard
The SWE-rebench leaderboard has been updated with a new multilingual slice that evaluates software engineering tasks across five programming languages: Go, Java, Python, Rust, and TypeScript. The update includes perform…
-
SWE-rebench leaderboard adds Claude Opus 4.8, GLM-5.2, Gemini 3.5 Flash
The SWE-rebench leaderboard has been updated with new models and improved UI, making it easier to compare AI performance on coding tasks. Notable additions include Claude Opus 4.8 xhigh, GLM-5.2, and Gemini 3.5 Flash, a…
-
2-bit GGUF models achieve 63% SWE-rebench pass rate with calibration
A new method has been developed to calibrate 2-bit quantized language models, specifically GGUF formats under 10GB, for agentic coding tasks. These calibrated models, such as Qwopus3.6-27B-Coder, achieve over 60% pass r…
-
SWE-rebench leaderboard adds 110 new Python tasks for AI models
The SWE-rebench leaderboard has been updated with 110 new Python tasks from GitHub PRs spanning March, April, and May. This update focuses on evaluating models' ability to read real issues, edit code, and pass test suit…