PulseAugur
EN
LIVE 08:22:29

New benchmark reveals widespread refusal failures in AI agents

Researchers have introduced HopRefusalBench, a new benchmark designed to evaluate how well search-augmented large language model agents handle unanswerable questions in multi-hop reasoning scenarios. The benchmark comprises 889 questions constructed from entity paths, covering three causes of unanswerability and three topologies of reasoning. Across ten proprietary and open-weight models tested, the best performer achieved only a 42.9% correct halting rate, indicating significant challenges in reliably refusing to answer when appropriate. AI

IMPACT Highlights critical limitations in current AI agents' ability to reliably refuse unanswerable queries, impacting their trustworthiness in complex reasoning tasks.

RANK_REASON The cluster contains a new academic paper introducing a novel benchmark for evaluating AI model capabilities. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New benchmark reveals widespread refusal failures in AI agents

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Jianan Xie, Xin Sun, Zhongqi Chen, Xing Zheng, Qiang Liu, Bowen Song ·

    HopRefusalBench: Diagnosing Refusal Failures in Search-Augmented Agents for Multi-Hop Reasoning

    arXiv:2608.01358v1 Announce Type: new Abstract: Search-augmented large language model agents are increasingly capable of solving knowledge-intensive tasks, but their behavior when a multi-hop question is fundamentally unanswerable remains poorly understood. Existing abstention be…