Researchers have introduced AISE-Bench, a new benchmark designed to evaluate large language models (LLMs) on their ability to seek information from academic knowledge graphs. This benchmark addresses limitations in existing tools by incorporating realistic user intents, complex multi-step API planning, and grounded answers with references. AISE-Bench includes over 1,100 question-answer pairs with detailed API execution trajectories and a comprehensive evaluation protocol. Initial testing showed that even advanced models like PLAY2PROMPT with Gemini-3-Pro achieved only moderate performance, highlighting significant challenges in API planning and execution for LLM agents. AI
IMPACT Establishes a new, challenging testbed for improving LLM agents' ability to interact with complex academic knowledge graphs.
RANK_REASON The cluster contains a new academic paper introducing a benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →