PulseAugur
EN
LIVE 13:44:01
ENTITY Agents Last Exam

Agents Last Exam

PulseAugur coverage of Agents Last Exam — every cluster mentioning Agents Last Exam across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
4
9 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
0
5 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
TIMELINE
  1. 2026-06-11 research_milestone GPT-5.5 outperformed Claude Fable 5 on the new Agents Last Exam benchmark. source
SENTIMENT · 30D

4 day(s) with sentiment data

RECENT · PAGE 1/1 · 9 TOTAL
  1. TOOL · CL_153338 ·

    AI Agents Last Exam Leaderboard Nearing Saturation by February

    The Agents Last Exam leaderboard is nearing saturation, with current benchmarks indicating it will be fully saturated by February of next year. This leaderboard tracks the performance of AI agents on various tasks, meas…

  2. SIGNIFICANT · CL_143284 ·

    GPT 5.6-Sol outperforms Claude Fable-5 on agentic tasks, but questions remain

    A new model, GPT 5.6-Sol, has reportedly surpassed Anthropic's Claude Fable-5 on long-horizon agentic tasks, achieving a score of 53.6 compared to Fable-5's 40.5 on the Agents' Last Exam. Despite this performance edge a…

  3. FRONTIER RELEASE · CL_131213 ·

    OpenAI launches GPT-5.6 family with Sol, Terra, and Luna tiers

    OpenAI has launched its new GPT-5.6 model family, featuring three tiers: Sol (flagship), Terra (balanced), and Luna (fastest and cheapest). These models offer varying price-performance points, with Sol being the most ca…

  4. RESEARCH · CL_115429 ·

    Meta AI hires AI safety expert Jianfeng Gao to lead Superintelligence Labs

    Meta AI has hired Dr. Yann LeCun's former colleague, Dr. Jianfeng Gao, to lead its Superintelligence Labs. Gao, previously a professor at UC Berkeley and co-founder of Virtue AI, will focus on AI safety and security eff…

  5. TOOL · CL_87018 ·

    New Benchmark Shows GPT 5.5 Outperforming Claude Fable 5 on Real-World Tasks

    A new benchmark called Agents' Last Exam (ALE), developed by researchers from UC Berkeley and other institutions, has revealed surprising results in AI agent performance. In the most challenging tasks, leading models li…

  6. RESEARCH · CL_85769 ·

    GPT-5.5 Outperforms Claude Fable 5 on New AI Agent Benchmark

    OpenAI's GPT-5.5 has reportedly outperformed Anthropic's Claude Fable 5 on the new Agents' Last Exam (ALE) benchmark. This benchmark, developed by UC Berkeley, evaluates AI agents' ability to perform complex, multi-step…

  7. SIGNIFICANT · CL_85182 ·

    GPT-5.5 surpasses Claude Fable 5 on new AI agent benchmark

    OpenAI's GPT-5.5 has outperformed Anthropic's Claude Fable 5 on a new AI benchmark called Agents Last Exam (ALE). This benchmark, developed by Berkeley RDI with input from over 300 experts, tests autonomous AI agents. T…

  8. TOOL · CL_72654 ·

    New benchmark tests AI agents on real-world economic tasks

    A new benchmark called Agents' Last Exam (ALE) has been introduced to evaluate AI agents on long-horizon, economically valuable tasks in real-world professional domains. Developed with over 250 industry experts, ALE cov…

  9. TOOL · CL_81353 ·

    New benchmark reveals AI agents pass only 2.6% of real-world tasks

    A new benchmark called Agents' Last Exam (ALE) has been introduced to evaluate AI agents on complex, real-world tasks relevant to professional industries. Developed with over 250 industry experts, ALE encompasses over 1…