PulseAugur
EN
LIVE 07:05:01
ENTITY Agents Last Exam

Agents Last Exam

PulseAugur coverage of Agents Last Exam — every cluster mentioning Agents Last Exam across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
4
8 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
0
0 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
TIMELINE
  1. 2026-06-11 research_milestone GPT-5.5 outperformed Claude Fable 5 on the new Agents Last Exam benchmark. source
SENTIMENT · 30D

4 day(s) with sentiment data

RECENT · PAGE 1/1 · 13 TOTAL
  1. FRONTIER RELEASE · CL_234775 ·

    OpenAI unveils GPT-6 Astra, claiming new state-of-the-art across benchmarks · 6 sources tracked

    OpenAI has announced GPT-6 Astra, its latest state-of-the-art model. The company claims Astra excels across a wide range of benchmarks, including FrontierMath Tier 4, ARC-AGI 3, and TerminalBench-4.0, demonstrating sign…

  2. SIGNIFICANT · CL_230221 ·

    DeepSeek V4 multimodal model weights released for inspection

    DeepSeek has released the weights and reference code for its V4 multimodal model, allowing researchers to examine its visual processing capabilities. Unlike simple image-to-text additions, V4 integrates visual tokens di…

  3. SIGNIFICANT · CL_210098 ·

    Qwen3.8-27B open-source model tops leaderboards with advanced agency

    Alibaba's Qwen team has open-sourced Qwen3.8-27B, a 27-billion parameter model that has achieved top rankings on several benchmarks, including SWE Bench Pro and OSWorld. This model demonstrates significant advancements …

  4. SIGNIFICANT · CL_203047 ·

    OpenAI's GPT-5.6 Sol excels at code generation but struggles with database population

    OpenAI has released GPT-5.6 Sol, which demonstrates significant improvements in coding tasks and token efficiency, outperforming previous models like Claude Opus 4.8 in benchmark tests. However, the model struggles with…

  5. TOOL · CL_153338 ·

    AI Agents Last Exam Leaderboard Nearing Saturation by February

    The Agents Last Exam leaderboard is nearing saturation, with current benchmarks indicating it will be fully saturated by February of next year. This leaderboard tracks the performance of AI agents on various tasks, meas…

  6. SIGNIFICANT · CL_143284 ·

    GPT 5.6-Sol outperforms Claude Fable-5 on agentic tasks, but questions remain

    A new model, GPT 5.6-Sol, has reportedly surpassed Anthropic's Claude Fable-5 on long-horizon agentic tasks, achieving a score of 53.6 compared to Fable-5's 40.5 on the Agents' Last Exam. Despite this performance edge a…

  7. FRONTIER RELEASE · CL_131213 ·

    OpenAI launches GPT-5.6 family with Sol, Terra, and Luna tiers

    OpenAI has launched its new GPT-5.6 model family, featuring three tiers: Sol (flagship), Terra (balanced), and Luna (fastest and cheapest). These models offer varying price-performance points, with Sol being the most ca…

  8. RESEARCH · CL_115429 ·

    Meta AI hires AI safety expert Jianfeng Gao to lead Superintelligence Labs

    Meta AI has hired Dr. Yann LeCun's former colleague, Dr. Jianfeng Gao, to lead its Superintelligence Labs. Gao, previously a professor at UC Berkeley and co-founder of Virtue AI, will focus on AI safety and security eff…

  9. TOOL · CL_87018 ·

    New Benchmark Shows GPT 5.5 Outperforming Claude Fable 5 on Real-World Tasks

    A new benchmark called Agents' Last Exam (ALE), developed by researchers from UC Berkeley and other institutions, has revealed surprising results in AI agent performance. In the most challenging tasks, leading models li…

  10. RESEARCH · CL_85769 ·

    GPT-5.5 Outperforms Claude Fable 5 on New AI Agent Benchmark

    OpenAI's GPT-5.5 has reportedly outperformed Anthropic's Claude Fable 5 on the new Agents' Last Exam (ALE) benchmark. This benchmark, developed by UC Berkeley, evaluates AI agents' ability to perform complex, multi-step…

  11. SIGNIFICANT · CL_85182 ·

    GPT-5.5 surpasses Claude Fable 5 on new AI agent benchmark

    OpenAI's GPT-5.5 has outperformed Anthropic's Claude Fable 5 on a new AI benchmark called Agents Last Exam (ALE). This benchmark, developed by Berkeley RDI with input from over 300 experts, tests autonomous AI agents. T…

  12. TOOL · CL_72654 ·

    New benchmark tests AI agents on real-world economic tasks

    A new benchmark called Agents' Last Exam (ALE) has been introduced to evaluate AI agents on long-horizon, economically valuable tasks in real-world professional domains. Developed with over 250 industry experts, ALE cov…

  13. TOOL · CL_81353 ·

    New benchmark reveals AI agents pass only 2.6% of real-world tasks

    A new benchmark called Agents' Last Exam (ALE) has been introduced to evaluate AI agents on complex, real-world tasks relevant to professional industries. Developed with over 250 industry experts, ALE encompasses over 1…