Researchers have introduced CatchBench, a novel benchmark designed to evaluate the ability of AI agents to detect their own failures. Unlike previous benchmarks, CatchBench assesses agents across three distinct information states: pre-run declared configurations, live traces during execution, and post-run finished traces. The benchmark includes seven task contracts with specific metrics, rather than a single leaderboard, to accommodate the different questions each state allows. Initial evaluations involved 72 entrants, including various LLM judges from major model families, across numerous configurations and runs, with many participants failing to achieve clear orderings. AI
IMPACT This benchmark could lead to more robust AI agents by improving their ability to self-diagnose and correct errors.
RANK_REASON The cluster contains a research paper introducing a new benchmark for evaluating AI agents. [lever_c_demoted from research: ic=1 ai=1.0]
Read on arXiv cs.MA (Multiagent) →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →