A new benchmark called CorporateBench (CB) has been introduced to evaluate Large Language Models (LLMs) on their ability to answer complex questions from large, temporally evolving enterprise document collections. Developed to address the limitations of existing benchmarks, CB includes over 230,000 documents and assesses LLMs across information extraction and knowledge base querying. Initial evaluations of five LLMs showed a significant drop in performance as the input size approached realistic corporate scales, highlighting a critical gap in current LLM reasoning capabilities for enterprise communication. AI
IMPACT Highlights a critical gap in LLM performance for enterprise data, potentially driving development of models better suited for corporate knowledge management.
RANK_REASON The cluster describes a new academic paper introducing a novel benchmark for evaluating LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
Read on arXiv cs.IR (Information Retrieval) →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →