English(EN)Mitigating Scoring Bias in LLM-as-a-Judge via Random Number Generation
新框架和方法解决LLM裁判的偏见 · 跟踪4个来源
作者PulseAugur 编辑部·[7 个来源]·
研究人员正在开发新方法来解决大型语言模型(LLM)在用作文本质量评估裁判时的评分偏差问题。一种方法是指示LLM生成随机数,以识别和纠正潜在的数值偏差,与现有方法相比表现有所提高。另一项研究引入了JudgeArena,这是一个统一的框架,可以跨不同基准和模型标准化LLM裁判的评估,从而提高可重复性和透明度。此外,一个名为JudgeBiasBench的新基准系统地量化了LLM裁判中不同类型的偏差,而一种称为Chain-of-Models的技术使用第二个LLM来审计主要LLM裁判的推理过程,证明了其在抵抗特定认知偏差方面的鲁棒性有所提高。
AI
arXiv:2608.05353v1 Announce Type: new Abstract: LLM judges are often asked to extract criteria and evidence before choosing between candidate answers. This workflow assumes that the intermediate record preserves the information needed for a later verdict. For reasoning-capable mo…
arXiv:2608.05726v1 Announce Type: new Abstract: Large Language Models (LLMs) are often used as evaluators of text quality, known as LLM-as-a-Judge, which can outperform conventional automatic evaluation metrics that rely on reference texts. However, LLM evaluators tend to generat…
Large Language Models (LLMs) are often used as evaluators of text quality, known as LLM-as-a-Judge, which can outperform conventional automatic evaluation metrics that rely on reference texts. However, LLM evaluators tend to generate particular scores regardless of the context of…
arXiv cs.CL
TIER_1English(EN)·Erlis Lushtaku, Bora Kargi, Ali Elganzory, Fabio Ferreira, Alejandro R. Salamanca, Julia Kreutzer, David Salinas·
arXiv:2608.02620v1 Announce Type: new Abstract: LLM-as-a-judge evaluation has become a dominant paradigm for ranking language models, yet the ecosystem remains fragmented: most benchmarks ship their own code base, hardcode a specific closed-model judge, and support a single evalu…
arXiv:2603.08091v2 Announce Type: replace Abstract: Large language model (LLM)-based judges are widely adopted for automated evaluation and reward modeling, yet their judgments are often affected by judgment biases. Accurately evaluating these biases is essential for ensuring the…
arXiv:2607.28636v1 Announce Type: new Abstract: LLMs increasingly serve as automated judges, but their judgments remain vulnerable to cognitive biases. Existing mitigations mostly rely on prompt-driven debiasing, which is brittle across bias types, or human evaluation, which does…
<p><strong>LLM-as-judge uses a capable model to score or compare agent outputs against a rubric</strong> — filling the gap where quality is open-ended and human judgment doesn't scale.</p> <p><strong>Done well it approximates human judgment cheaply; done carelessly it produces co…