PulseAugur
EN
LIVE 22:08:46

GPT-5.5 and Claude Fable 5 fail new ActiveVision benchmark, scoring far below humans

A new benchmark called ActiveVision, designed to test repeated visual perception, has revealed significant limitations in advanced AI models. GPT-5.5 achieved only a 10.6% success rate, failing entirely on 11 out of 17 tasks. Claude Fable 5 performed even worse with a 3.5% score, starkly contrasting with human participants who averaged 96.1% accuracy. The study highlights that even high-tier models struggle with tasks requiring continuous visual processing and cannot compensate by generating their own code. AI

IMPACT Highlights current AI limitations in continuous visual perception tasks, indicating a need for new architectures beyond current reasoning and coding capabilities.

RANK_REASON New benchmark evaluation of existing frontier models. [lever_c_demoted from research: ic=1 ai=1.0]

Read on r/MachineLearning →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

GPT-5.5 and Claude Fable 5 fail new ActiveVision benchmark, scoring far below humans

COVERAGE [1]

  1. r/MachineLearning TIER_1 English(EN) · /u/Justgototheeffinmoon ·

    GPT-5.5 Scores 10.6% on ActiveVision, Humans Hit 96.1% [R]

    <!-- SC_OFF --><div class="md"><p>The interesting finding from a new [arXiv paper](<a href="https://arxiv.org/abs/2607.16165">https://arxiv.org/abs/2607.16165</a>) isn't that a frontier vision model failed a new benchmark, that happens weekly, but the specific shape of the failur…