A new benchmark, ESQ-Bench, has been developed to evaluate Natural Language to SQL (NL2SQL) models on enterprise database environments, which are more complex than typical academic benchmarks. The benchmark includes six populated schemas and 550 question-query pairs across three complexity tiers, using Oracle as the primary database. Initial tests show that Claude Sonnet 4.6 outperforms GPT-4o on ESQ-Bench, particularly at higher complexity tiers, while open-weight models like Llama 3.2 struggle significantly. AI
影响 This benchmark could drive improvements in enterprise-grade NL2SQL capabilities, pushing models to better handle complex database schemas and dialects.
排序理由 The item describes a new academic benchmark for evaluating NL2SQL models. [lever_c_demoted from research: ic=1 ai=1.0]
AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →