Researchers have introduced ESF-Bench, a new benchmark designed to evaluate the performance of large language models (LLMs) in challenging slot-filling scenarios relevant to enterprise applications. The benchmark comprises 810 multi-turn samples and 6530 slots across 8 domains, focusing on 57 difficult real-world enterprise deployment scenarios. Initial testing revealed significant limitations in current LLMs, with GPT-OSS-120b achieving only a 20.7% success rate in slot extraction. The dataset, taxonomy, and evaluation code have been made publicly available on GitHub to foster further research. AI
IMPACT Highlights critical gaps in LLM capabilities for real-world enterprise data extraction, potentially guiding future model development.
RANK_REASON The cluster describes a new academic benchmark for evaluating LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →