PulseAugur
中
实时 00:05:10
English(EN) Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline

文本到SQL管道中的LLM作为裁判显示出重大故障,Qwen3.6-27B提供改进

arXiv上的一篇新研究论文详细介绍了在生产文本到SQL管道中使用大型语言模型(LLM)作为裁判时出现的重大故障。研究发现,GPT-4o mini在用作裁判时,与人类标注者的一致性较低,由于一种称为GRADE-HALLUCINATION的机制,经常过度标记正确的SQL查询。用自托管的Qwen3.6-27B模型替换GPT-4o mini提高了其一致性,并提供了低得多的每次调用成本,而集成多个强大的裁判模型进一步提高了准确性和覆盖率。 AI

影响 强调了在生产系统中对LLM裁判进行严格审计的必要性,并确定了Qwen3.6-27B等具有成本效益的替代方案。

排序理由 该集群包含一篇研究论文,详细介绍了对特定应用程序中LLM故障的审计和修复。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

文本到SQL管道中的LLM作为裁判显示出重大故障,Qwen3.6-27B提供改进

报道来源 [1]

  1. arXiv cs.LG TIER_1 English(EN) · Haowei Liu, Hsin-Tai Wu, Yi Fang ·

    生产文本到SQL管道中审计和修复LLM-as-Judge的故障

    arXiv:2609.30290v1 Announce Type: cross Abstract: Production text-to-SQL pipelines often end with an LLM-as-judge whose agreement with human annotators has never actually been measured. When we checked ours, the deployed gpt-4o-mini judge agreed with two-author gold at only Cohen…