PulseAugur
EN
LIVE 00:05:07

LLM-as-Judge in Text-to-SQL Pipelines Shows Significant Failures, Qwen3.6-27B Offers Improvement

A new research paper from arXiv details significant failures in using large language models (LLMs) as judges in production text-to-SQL pipelines. The study found that GPT-4o mini, when used as a judge, had low agreement with human annotators, often over-flagging correct SQL queries due to a mechanism termed GRADE-HALLUCINATION. Replacing GPT-4o mini with a self-hosted Qwen3.6-27B model improved agreement and offered a much lower cost per call, while ensembling multiple strong judges further enhanced accuracy and coverage. AI

IMPACT Highlights the need for rigorous auditing of LLM judges in production systems and identifies cost-effective alternatives like Qwen3.6-27B.

RANK_REASON The cluster contains a research paper detailing an audit and repair of LLM failures in a specific application. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.LG →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

LLM-as-Judge in Text-to-SQL Pipelines Shows Significant Failures, Qwen3.6-27B Offers Improvement

COVERAGE [1]

  1. arXiv cs.LG TIER_1 English(EN) · Haowei Liu, Hsin-Tai Wu, Yi Fang ·

    Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline

    arXiv:2609.30290v1 Announce Type: cross Abstract: Production text-to-SQL pipelines often end with an LLM-as-judge whose agreement with human annotators has never actually been measured. When we checked ours, the deployed gpt-4o-mini judge agreed with two-author gold at only Cohen…