PulseAugur
EN
LIVE 09:32:57

New VAKRA benchmark reveals AI agent struggles with multi-hop reasoning

Researchers have introduced VAKRA, a new benchmark designed to evaluate the multi-hop reasoning capabilities of AI agents that interact with both structured APIs and document collections. The benchmark includes over 8,000 executable APIs across 62 domains and tasks of increasing difficulty, including policy constraints. Initial evaluations using a fixed ReAct harness revealed that even state-of-the-art models struggle with complex reasoning, with performance degrading significantly as reasoning depth increases and severe failures occurring when adhering to tool-use policies. AI

IMPACT Highlights significant limitations in current AI agent reasoning capabilities, particularly with complex API interactions and policy adherence.

RANK_REASON The cluster describes a new academic paper introducing a benchmark for evaluating AI agents. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New VAKRA benchmark reveals AI agent struggles with multi-hop reasoning

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Ankita Rajaram Naik, Anupama Murthi, Benjamin Elder, Siyu Huo, Raavi Gupta, Abhinav Jain, Praveen Venkateswaran, Abdulhamid Adebayo, Danish Contractor ·

    VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies

    arXiv:2608.12282v1 Announce Type: new Abstract: Agents deployed in enterprise settings must reason across structured APIs and document collections, yet existing benchmarks evaluate these capabilities in isolation. We introduce VAKRA (e\textbf{V}aluating \textbf{A}PI and \textbf{K…