PulseAugur
EN
LIVE 07:32:01

New AISE-Bench benchmark challenges LLMs on academic knowledge graph information seeking

Researchers have introduced AISE-Bench, a new benchmark designed to evaluate large language models (LLMs) on their ability to seek information from academic knowledge graphs. This benchmark addresses limitations in existing tools by incorporating realistic user intents, complex multi-step API planning, and grounded answers with references. AISE-Bench includes over 1,100 question-answer pairs with detailed API execution trajectories and a comprehensive evaluation protocol. Initial testing showed that even advanced models like PLAY2PROMPT with Gemini-3-Pro achieved only moderate performance, highlighting significant challenges in API planning and execution for LLM agents. AI

IMPACT Establishes a new, challenging testbed for improving LLM agents' ability to interact with complex academic knowledge graphs.

RANK_REASON The cluster contains a new academic paper introducing a benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New AISE-Bench benchmark challenges LLMs on academic knowledge graph information seeking

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Fanjin Zhang, Zhengyang Wang, Ruixuan Huang, Kefan Zhang, Amy Xin, Yuanchun Wang, Shu Zhao, Evgeny Kharlamov, Jie Tang, Juanzi Li ·

    AISE-Bench: A Full-Cycle Curated Benchmark for Information Seeking on Academic Knowledge Graphs

    arXiv:2607.20498v1 Announce Type: new Abstract: Large language models (LLMs) augmented with tools are emerging as autonomous agents capable of using Web engine, APIs, and code to solve complex, long-horizon tasks. Current tool-using benchmarks for information seeking on academic …