Researchers have introduced BIRD-History, a new benchmark designed to evaluate text-to-SQL systems' ability to leverage historical query logs for understanding underspecified natural language questions. The benchmark includes 1,393 tasks across 11 databases, with annotations to identify relevant knowledge in past SQL scripts. A proposed plug-in retriever extracts and reranks external knowledge from these historical logs, demonstrating consistent performance improvements when integrated with existing text-to-SQL systems. AI
IMPACT This benchmark could improve the accuracy of text-to-SQL systems by enabling them to better understand context from historical data.
RANK_REASON The item describes a new benchmark and associated methodology for evaluating AI systems, published on arXiv. [lever_c_demoted from research: ic=1 ai=1.0]
- BIRD-History
- Few-shot learning
- Hugging Face
- large language model
- natural language
- SQL
- SQL query logs
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →