PulseAugur
实时 18:55:38
English(EN) Does the Judge Prefer English? Evaluating Language-Switching Invariance in LLM-as-a-Judge

研究发现:大型语言模型法官存在语言偏见

一项新研究评估了大型语言模型(LLMs)在评估响应时作为法官的功能,发现语言偏好会显著影响其判断。该研究引入了Judge-LS协议,该协议使用英语、中文以及响应对的语言切换变体来测试LLM法官。结果显示,与英语相比,中文和语言切换的呈现方式在10.7%至14.4%的案例中导致了偏好翻转,所有测试的法官在英语中的表现最佳。然而,该研究并未发现当翻译等效时存在系统性的偏向英语的偏见,因为大多数此类探测被判断为平局,而非平局的决定有时偏向中文。 AI

影响 揭示了LLM评估指标中潜在的偏见,强调了对更鲁棒、语言不变性评估方法的需求。

排序理由 该集群包含一篇详细介绍LLM新评估协议的学术论文。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

研究发现:大型语言模型法官存在语言偏见

报道来源 [2]

  1. arXiv cs.CL TIER_1 English(EN) · Shaojie Yin ·

    法官偏爱英语吗?评估LLM作为法官时的语言切换不变性

    arXiv:2606.14278v1 Announce Type: new Abstract: Large language models (LLMs) are now widely used as automatic judges for open-ended instruction-following evaluation. This practice is convenient, scalable, and often more semantically aware than reference-based metrics, but it also…

  2. arXiv cs.CL TIER_1 English(EN) · Shaojie Yin ·

    法官偏爱英语吗?评估LLM作为法官时的语言切换不变性

    Large language models (LLMs) are now widely used as automatic judges for open-ended instruction-following evaluation. This practice is convenient, scalable, and often more semantically aware than reference-based metrics, but it also introduces a new reliability question: does a j…