Anthropic's Opus 5 model has achieved a new state-of-the-art performance on the ProgramBench benchmark, successfully solving 9 out of 200 task instances. This represents a more than fourfold improvement compared to the previous best, GPT Sol. The evaluation of Opus 5 on this benchmark was notably expensive, costing $10,000. AI
IMPACT Sets a new benchmark for agentic program synthesis, potentially influencing future model development in code generation.
RANK_REASON Research milestone on a specific benchmark. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →