Researchers have introduced TIDE-Bench, a new benchmark designed to evaluate large language models (LLMs) on conversational text-to-SQL tasks, specifically focusing on chain ambiguity and intent drift. This benchmark, built upon 514 anchor SQLs from the BIRD dataset, includes 1,542 samples and introduces metrics for identifying chain ambiguities and resolving intent drifts, going beyond simple execution accuracy. Evaluations of 12 LLMs using TIDE-Bench revealed significant challenges in chain identification and a notable gap in recognizing and resolving user intent drift. AI
IMPACT This benchmark could lead to more robust conversational AI systems capable of handling complex user queries and revisions.
RANK_REASON The cluster contains a research paper introducing a new benchmark for evaluating LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →