PulseAugur
EN
LIVE 17:34:54

Anthropic's Opus 5 sets new SoTA on ProgramBench, beating GPT Sol

Anthropic's Opus 5 model has achieved a new state-of-the-art performance on the ProgramBench benchmark, successfully solving 9 out of 200 task instances. This represents a more than fourfold improvement compared to the previous best, GPT Sol. The evaluation of Opus 5 on this benchmark was notably expensive, costing $10,000. AI

IMPACT Sets a new benchmark for agentic program synthesis, potentially influencing future model development in code generation.

RANK_REASON Research milestone on a specific benchmark. [lever_c_demoted from research: ic=1 ai=1.0]

Read on r/Anthropic →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Anthropic's Opus 5 sets new SoTA on ProgramBench, beating GPT Sol

COVERAGE [1]

  1. r/Anthropic TIER_1 English(EN) · /u/klieret ·

    Opus 5 setting new SoTA on ProgramBench, solves 9/200 instances (>4x than GPT Sol)

    <table> <tr><td> <a href="https://www.reddit.com/r/Anthropic/comments/1vnfhla/opus_5_setting_new_sota_on_programbench_solves/"> <img alt="Opus 5 setting new SoTA on ProgramBench, solves 9/200 instances (&gt;4x than GPT Sol)" src="https://preview.redd.it/lccofwq346jh1.png?width=14…