PulseAugur
EN
LIVE 21:08:40

Claude Opus 4.5 excels on Senior SWE-bench, ranking second overall

A new benchmark, Senior SWE-bench, has been developed to evaluate AI models on tasks typically performed by senior software engineers. In initial tests, Claude Opus 4.5 demonstrated strong performance, ranking second overall and first for bug investigation tasks. The benchmark includes 100 tasks, with 50 reserved to prevent data contamination. AI

IMPACT Establishes a new benchmark for evaluating AI models on complex engineering tasks, potentially driving future model development.

RANK_REASON New benchmark evaluation of an AI model. [lever_c_demoted from research: ic=1 ai=1.0]

Read on r/Anthropic →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Claude Opus 4.5 excels on Senior SWE-bench, ranking second overall

COVERAGE [1]

  1. r/Anthropic TIER_1 English(EN) · /u/PubliusAu ·

    Opus 5 nearly matches Fable 5 on senior engineering tasks using ~32% of the output tokens on average

    <table> <tr><td> <a href="https://www.reddit.com/r/Anthropic/comments/1v5jo21/opus_5_nearly_matches_fable_5_on_senior/"> <img alt="Opus 5 nearly matches Fable 5 on senior engineering tasks using ~32% of the output tokens on average" src="https://preview.redd.it/zq1d0bqc08fh1.png?…