The text-to-SQL benchmark landscape, dominated by BIRD and Spider, often uses schemas that do not reflect real-world user applications. The author introduces 'persona-bench,' a new benchmark designed with schemas and queries representative of typical user interactions with products like nlqdb. This new benchmark, using the same scoring mechanism as BIRD and Spider, shows a significantly higher accuracy for their system, highlighting the importance of using relevant benchmarks for evaluating text-to-SQL performance in practical scenarios. AI
IMPACT Highlights the need for more realistic benchmarks in text-to-SQL to better reflect real-world application performance.
RANK_REASON Introduces a new benchmark for evaluating text-to-SQL models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →