Researchers have developed PAIChecker, a multi-agent system designed to identify and correct misalignments between pull requests (PRs) and their associated issues in benchmarks used to evaluate large language models (LLMs). A study of SWE-bench Verified instances revealed that 13.6% of PR-issue pairings exhibit misalignment across various patterns. PAIChecker employs a three-phase approach combining pattern identification, label synthesis, and code-level validation to ensure more accurate and generalizable benchmark construction. Experiments demonstrated PAIChecker's superior performance on SWE-Gym and SWE-bench Multilingual datasets, achieving up to 92.12% binary accuracy. AI
IMPACT Improves the reliability of benchmarks used to evaluate LLM capabilities in software development tasks.
RANK_REASON The cluster describes a new research paper introducing a novel system for benchmark validation. [lever_c_demoted from research: ic=1 ai=1.0]
- edition or translation
- Manyi Wang
- PAIChecker
- SWE-bench
- SWE-Bench Multilingual
- SWE-bench Verified
- SWE-Gym
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →