A new research paper from arXiv details significant failures in using large language models (LLMs) as judges in production text-to-SQL pipelines. The study found that GPT-4o mini, when used as a judge, had low agreement with human annotators, often over-flagging correct SQL queries due to a mechanism termed GRADE-HALLUCINATION. Replacing GPT-4o mini with a self-hosted Qwen3.6-27B model improved agreement and offered a much lower cost per call, while ensembling multiple strong judges further enhanced accuracy and coverage. AI
IMPACT Highlights the need for rigorous auditing of LLM judges in production systems and identifies cost-effective alternatives like Qwen3.6-27B.
RANK_REASON The cluster contains a research paper detailing an audit and repair of LLM failures in a specific application. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →