A new benchmark called ARCS has been developed to address the challenge of ambiguity in text-to-SQL systems. This benchmark features naturally occurring ambiguities over real-world databases, complete with annotations for valid ambiguity points, interpretations, and corresponding SQL queries. Initial experiments with ARCS show that current text-to-SQL models struggle with ambiguity, with the GPT-6 Sol model achieving only 51% execution accuracy and open-source models performing below 27%. The proposed structured disambiguation paradigm aims to resolve these ambiguities through explicit, constrained interactions rather than conversational clarification. AI
IMPACT Highlights the significant challenges in developing robust text-to-SQL systems that can handle real-world ambiguity, potentially guiding future research in this area.
RANK_REASON The cluster describes a new academic benchmark and associated research paper. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →