PulseAugur
EN
LIVE 20:25:38

METR AI time horizons graph riddled with severe errors, analysis finds

A recent analysis by Nathan Witkin, a research writer at NYU Stern’s Tech and Society Lab, has identified numerous severe errors in the widely cited METR AI time horizons graph. These flaws include fabricated human baseline data, incentivizing benchmarkers to take longer by paying them hourly, a biased sample of human testers, and potential test-training data contamination. Witkin argues that the graph's significant inaccuracies render it unreliable for drawing meaningful conclusions about AI capabilities and their impact on tasks like software development. AI

IMPACT Critiques of widely cited AI capability graphs highlight the need for rigorous scientific standards and can influence how AI progress is perceived.

RANK_REASON The cluster discusses a critique of a previously published graph, rather than a new release or research finding.

Read on r/MachineLearning →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

COVERAGE [1]

  1. r/MachineLearning TIER_1 · /u/common_yarrow ·

    The famous METR AI time horizons graph contains numerous severe errors [D]

    <!-- SC_OFF --><div class="md"><p>Nathan Witkin, a research writer at NYU Stern’s Tech and Society Lab, <a href="https://www.transformernews.ai/p/against-the-metr-graph-coding-capabilities-software-jobs-task-ai">writes</a> damningly about the famous METR AI time horizons graph in…