Researchers have introduced CatchBench, a new benchmark designed to evaluate how effectively AI agent failures can be detected. Unlike previous benchmarks, CatchBench assesses agents across three distinct information states: pre-run declared configurations, live trace prefixes, and complete post-run traces. The benchmark includes seven task contracts with specific metrics, rather than a single leaderboard, to accommodate the different questions each state allows. It evaluated 72 entrants, including eleven LLM judges from nine model families, across numerous configurations and runs, finding that most participants did not establish clear orderings. AI
IMPACT Introduces a novel evaluation framework for AI agents, potentially improving the reliability and auditability of AI systems.
RANK_REASON The cluster contains a research paper detailing a new benchmark for AI agents.
Read on arXiv cs.MA (Multiagent) →
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →