Researchers have introduced VAKRA, a new benchmark designed to evaluate the multi-hop reasoning capabilities of AI agents that interact with both structured APIs and document collections. The benchmark includes over 8,000 executable APIs across 62 domains and tasks of increasing difficulty, including policy constraints. Initial evaluations using a fixed ReAct harness revealed that even state-of-the-art models struggle with complex reasoning, with performance degrading significantly as reasoning depth increases and severe failures occurring when adhering to tool-use policies. AI
IMPACT Highlights significant limitations in current AI agent reasoning capabilities, particularly with complex API interactions and policy adherence.
RANK_REASON The cluster describes a new academic paper introducing a benchmark for evaluating AI agents. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- Ankita Rajaram Naik
- arXiv
- CatalyzeX Code Finder for Papers
- DagsHub
- Gotit.pub
- Hugging Face
- ReAct
- ScienceCast
- VAKRA
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →