Researchers have introduced XL-DocBench, a new benchmark designed to evaluate the capabilities of large language models in understanding extremely long documents. This benchmark features over 1,500 questions derived from professional domains like annual reports and clinical guidelines, with documents extending up to 2,303 pages. XL-DocBench requires models to not only locate information but also to synthesize evidence from multiple pages and documents, and interpret tables, charts, and figures, with a significant portion of questions requiring multi-document evidence. Initial results indicate that current LLMs still face challenges in handling such extensive contexts and complex reasoning tasks. AI
IMPACT This benchmark will drive development of LLMs capable of processing and reasoning over extensive professional documents, crucial for applications in compliance, finance, and healthcare.
RANK_REASON The cluster contains a research paper introducing a new benchmark for evaluating LLM capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →