A new benchmark, ESQ-Bench, has been developed to evaluate Natural Language to SQL (NL2SQL) models on enterprise database environments, which are more complex than typical academic benchmarks. The benchmark includes six populated schemas and 550 question-query pairs across three complexity tiers, using Oracle as the primary database. Initial tests show that Claude Sonnet 4.6 outperforms GPT-4o on ESQ-Bench, particularly at higher complexity tiers, while open-weight models like Llama 3.2 struggle significantly. AI
IMPACT This benchmark could drive improvements in enterprise-grade NL2SQL capabilities, pushing models to better handle complex database schemas and dialects.
RANK_REASON The item describes a new academic benchmark for evaluating NL2SQL models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →