Researchers have introduced ToolRobustBench, a new benchmark designed to evaluate and diagnose failures in tool-calling agents, which are LLM systems that use external tools to complete tasks. The benchmark systematically introduces perturbations across four categories—tool-interface, user-intent, tool-output/observation, and runtime-environment—to pinpoint specific failure points within the tool-use pipeline. Experiments involving multiple LLMs and tools revealed significant robustness degradation, with tool-output/observation perturbations identified as a major bottleneck, highlighting the need for more robust tool-calling capabilities. AI
影响 Enhances evaluation of LLM agent robustness, guiding development towards more reliable tool integration.
排序理由 The cluster contains a research paper introducing a new benchmark for evaluating LLM tool-calling capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- LLMs
- runtime-environment perturbations
- tool calling
- tool-interface perturbations
- tool-output/observation perturbations
- ToolRobustBench
- tool-use pipeline
- user-intent perturbations
AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →