OpenAI's GPT-5.5 has reportedly outperformed Anthropic's Claude Fable 5 on the new Agents' Last Exam (ALE) benchmark. This benchmark, developed by UC Berkeley, evaluates AI agents' ability to perform complex, multi-step tasks autonomously. GPT-5.5 achieved a score of 94%, while Claude Fable 5 scored 86% on the challenging test. AI
IMPACT This benchmark result suggests GPT-5.5 may have superior autonomous task execution capabilities compared to Claude Fable 5.
RANK_REASON The cluster reports on a new benchmark result for AI models, which falls under research.
Read on Mastodon — sigmoid.social →
AI-generated summary · Google Gemini · from 4 sources. How we write summaries →