A new benchmark called Legal Research Bench (LRB) has been developed to measure the end-to-end reliability of AI agents in performing complex legal research tasks. The benchmark consists of 413 open-ended questions created by legal experts, along with gold answers and a grading rubric. When tested, even the top-performing model, Claude Opus 4.8, achieved only 42.9% accuracy, indicating that current AI agents are far from reliable for critical legal workflows. Performance varied by task, with questions requiring reconciliation of conflicting authorities proving particularly challenging. AI
IMPACT Highlights significant gaps in AI reliability for critical, high-stakes applications like legal research, indicating a need for further development in agent reasoning and fact-verification.
RANK_REASON The cluster contains a research paper introducing a new benchmark for evaluating AI capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →