PulseAugur
EN
LIVE 09:40:51

New ESF-Bench highlights LLM limitations in enterprise slot-filling

Researchers have introduced ESF-Bench, a new benchmark designed to evaluate the performance of large language models (LLMs) in challenging slot-filling scenarios relevant to enterprise applications. The benchmark comprises 810 multi-turn samples and 6530 slots across 8 domains, focusing on 57 difficult real-world enterprise deployment scenarios. Initial testing revealed significant limitations in current LLMs, with GPT-OSS-120b achieving only a 20.7% success rate in slot extraction. The dataset, taxonomy, and evaluation code have been made publicly available on GitHub to foster further research. AI

IMPACT Highlights critical gaps in LLM capabilities for real-world enterprise data extraction, potentially guiding future model development.

RANK_REASON The cluster describes a new academic benchmark for evaluating LLMs. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New ESF-Bench highlights LLM limitations in enterprise slot-filling

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Toby Liang, Gopal Sarda, Sagar Davasam, Vikas Yadav ·

    ESF-Bench: Benchmarking Challenging Slot-Filling Scenarios for Real-World Enterprise Applications

    arXiv:2607.23326v1 Announce Type: new Abstract: The rapid rise of large language models (LLMs) has driven transformative adoption across enterprises. However, deploying these models in real-world settings presents unique challenges due to complex system constraints and unexpected…