Researchers have introduced SWE-Bench ProMax, a new benchmark designed to evaluate AI coding agents on complex, large-scale code refactoring tasks. This benchmark addresses limitations in existing evaluations, such as flawed tests and the ability of models to reproduce training data. SWE-Bench ProMax features 170 expert-curated instances across seven programming languages, with an average of 11.4 modified files and 261.6 lines of code per instance. Initial experiments show that even advanced AI models achieve only a 41.2% resolution rate, indicating that SWE-Bench ProMax presents a significant and currently challenging test for AI coding capabilities. AI
IMPACT This benchmark will push the development of more capable AI coding agents by providing a more rigorous evaluation of their ability to handle complex, multi-file code refactoring.
RANK_REASON The cluster describes a new academic benchmark for evaluating AI coding agents.
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →