PulseAugur
EN
LIVE 16:53:15

New EntSQL benchmark tests Text-to-SQL in enterprise knowledge

Researchers have introduced EntSQL, a new benchmark designed to evaluate Text-to-SQL capabilities in enterprise settings. Unlike previous benchmarks, EntSQL focuses on grounding SQL generation in long-context, proprietary business documents. The benchmark includes 1,066 aligned Chinese-English examples across five business domains, many of which require knowledge beyond the immediate question and schema. Current systems struggle with this task, with the best performing model achieving only 15.9% accuracy on English inputs when provided with long-form documents. AI

IMPACT Highlights the challenge of applying LLMs to enterprise-specific data, potentially driving development of more context-aware Text-to-SQL systems.

RANK_REASON The cluster describes a new academic benchmark for evaluating AI capabilities.

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

New EntSQL benchmark tests Text-to-SQL in enterprise knowledge

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
The cluster describes a new academic benchmark for evaluating AI capabilities.
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
116 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

Full methodology in our editorial standards.

COVERAGE [2]

  1. arXiv cs.CL TIER_1 English(EN) · Chengxi Liao, Tao Xu, Zulong Chen, Chuanfei Xu, Yiyan Wang, Xinyun Wang, Yanlong Zhang, Xiaojun Chen, Zhibo Yang, Zeyi Wen ·

    EntSQL: A Benchmark for Grounding Text-to-SQL in Long-Context Enterprise Knowledge

    arXiv:2606.03363v1 Announce Type: new Abstract: Text-to-SQL enables natural language access to databases, and recent LLMs have substantially advanced its capabilities. Existing benchmarks such as Spider, BIRD, and Spider~2.0 evaluate schema generalization, large-scale databases, …

  2. arXiv cs.CL TIER_1 English(EN) · Zeyi Wen ·

    EntSQL: A Benchmark for Grounding Text-to-SQL in Long-Context Enterprise Knowledge

    Text-to-SQL enables natural language access to databases, and recent LLMs have substantially advanced its capabilities. Existing benchmarks such as Spider, BIRD, and Spider~2.0 evaluate schema generalization, large-scale databases, and realistic workflows, but largely overlook en…