A new benchmark called DAYJOB has been developed to evaluate AI agents on long-horizon professional tasks, particularly in healthcare and finance. The benchmark consists of 130 tasks, each requiring an average of 13.6 to 16.6 hours for a human professional to complete. Even advanced models like Claude Opus 5.5 struggle, passing only around 24% of healthcare and finance tasks, indicating significant challenges for AI in complex, multi-step professional work. The researchers found that agents often accept flawed premises and propagate errors throughout their analyses. AI
IMPACT Highlights significant limitations in current AI capabilities for complex, multi-step professional tasks, indicating a need for improved reasoning and error-handling.
RANK_REASON The cluster describes a new benchmark for AI evaluation, presented in an arXiv paper. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →