Agents Last Exam
PulseAugur coverage of Agents Last Exam — every cluster mentioning Agents Last Exam across labs, papers, and developer communities, ranked by signal.
- 2026-06-11 research_milestone GPT-5.5 outperformed Claude Fable 5 on the new Agents Last Exam benchmark. source
4 day(s) with sentiment data
-
OpenAI unveils GPT-6 Astra, claiming new state-of-the-art across benchmarks · 6 sources tracked
OpenAI has announced GPT-6 Astra, its latest state-of-the-art model. The company claims Astra excels across a wide range of benchmarks, including FrontierMath Tier 4, ARC-AGI 3, and TerminalBench-4.0, demonstrating sign…
-
DeepSeek V4 multimodal model weights released for inspection
DeepSeek has released the weights and reference code for its V4 multimodal model, allowing researchers to examine its visual processing capabilities. Unlike simple image-to-text additions, V4 integrates visual tokens di…
-
Qwen3.8-27B open-source model tops leaderboards with advanced agency
Alibaba's Qwen team has open-sourced Qwen3.8-27B, a 27-billion parameter model that has achieved top rankings on several benchmarks, including SWE Bench Pro and OSWorld. This model demonstrates significant advancements …
-
OpenAI's GPT-5.6 Sol excels at code generation but struggles with database population
OpenAI has released GPT-5.6 Sol, which demonstrates significant improvements in coding tasks and token efficiency, outperforming previous models like Claude Opus 4.8 in benchmark tests. However, the model struggles with…
-
AI Agents Last Exam Leaderboard Nearing Saturation by February
The Agents Last Exam leaderboard is nearing saturation, with current benchmarks indicating it will be fully saturated by February of next year. This leaderboard tracks the performance of AI agents on various tasks, meas…
-
GPT 5.6-Sol outperforms Claude Fable-5 on agentic tasks, but questions remain
A new model, GPT 5.6-Sol, has reportedly surpassed Anthropic's Claude Fable-5 on long-horizon agentic tasks, achieving a score of 53.6 compared to Fable-5's 40.5 on the Agents' Last Exam. Despite this performance edge a…
-
OpenAI launches GPT-5.6 family with Sol, Terra, and Luna tiers
OpenAI has launched its new GPT-5.6 model family, featuring three tiers: Sol (flagship), Terra (balanced), and Luna (fastest and cheapest). These models offer varying price-performance points, with Sol being the most ca…
-
Meta AI hires AI safety expert Jianfeng Gao to lead Superintelligence Labs
Meta AI has hired Dr. Yann LeCun's former colleague, Dr. Jianfeng Gao, to lead its Superintelligence Labs. Gao, previously a professor at UC Berkeley and co-founder of Virtue AI, will focus on AI safety and security eff…
-
New Benchmark Shows GPT 5.5 Outperforming Claude Fable 5 on Real-World Tasks
A new benchmark called Agents' Last Exam (ALE), developed by researchers from UC Berkeley and other institutions, has revealed surprising results in AI agent performance. In the most challenging tasks, leading models li…
-
GPT-5.5 Outperforms Claude Fable 5 on New AI Agent Benchmark
OpenAI's GPT-5.5 has reportedly outperformed Anthropic's Claude Fable 5 on the new Agents' Last Exam (ALE) benchmark. This benchmark, developed by UC Berkeley, evaluates AI agents' ability to perform complex, multi-step…
-
GPT-5.5 surpasses Claude Fable 5 on new AI agent benchmark
OpenAI's GPT-5.5 has outperformed Anthropic's Claude Fable 5 on a new AI benchmark called Agents Last Exam (ALE). This benchmark, developed by Berkeley RDI with input from over 300 experts, tests autonomous AI agents. T…
-
New benchmark tests AI agents on real-world economic tasks
A new benchmark called Agents' Last Exam (ALE) has been introduced to evaluate AI agents on long-horizon, economically valuable tasks in real-world professional domains. Developed with over 250 industry experts, ALE cov…
-
New benchmark reveals AI agents pass only 2.6% of real-world tasks
A new benchmark called Agents' Last Exam (ALE) has been introduced to evaluate AI agents on complex, real-world tasks relevant to professional industries. Developed with over 250 industry experts, ALE encompasses over 1…