A new benchmark called MirrorCode has been developed to evaluate AI models on large-scale, long-horizon software engineering tasks. The benchmark involves reimplementing entire programs from scratch, with one task costing $2,600 and taking 19 days for an AI to complete. Notably, Claude Opus 4.7 successfully reimplemented a bioinformatics toolkit named gotree, a task estimated to take a human engineer weeks, in just 14 hours for $251. While the benchmark aims to be cheat-resistant, the possibility of data contamination from pre-training on open-source code remains a consideration. AI
IMPACT This benchmark could accelerate the development of AI agents capable of complex, long-horizon software engineering tasks.
RANK_REASON The cluster describes a new benchmark for evaluating AI capabilities in software engineering, including details of its construction and initial results.
Read on Mastodon — mastodon.social →
AI-generated summary · Google Gemini · from 5 sources. How we write summaries →