A new benchmark, Senior SWE-bench, has been developed to evaluate AI models on tasks typically performed by senior software engineers. In initial tests, Claude Opus 4.5 demonstrated strong performance, ranking second overall and first for bug investigation tasks. The benchmark includes 100 tasks, with 50 reserved to prevent data contamination. AI
IMPACT Establishes a new benchmark for evaluating AI models on complex engineering tasks, potentially driving future model development.
RANK_REASON New benchmark evaluation of an AI model. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →