PulseAugur
EN
LIVE 09:00:01

LLM judges show bidirectional bias based on self-labeling

A new research paper explores the phenomenon of self-preference bias in Large Language Models (LLMs) when they are used as judges for evaluating outputs. The study found that when LLM judges were unaware of the source of the selections, self-preference largely disappeared after controlling for selection quality and evaluator severity. However, when labels indicating self- or other-authorship were present, LLMs exhibited bidirectional bias, inflating scores for self-labeled selections and deflating scores for other-labeled ones, even without explicit model identification. AI

IMPACT Highlights potential biases in LLM evaluation systems, impacting the reliability of AI-generated content assessment.

RANK_REASON Research paper published on arXiv detailing findings about LLM bias. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

LLM judges show bidirectional bias based on self-labeling

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Songeun Chae, Min Kim, Donghoon Jung, Seojin Choi, Seohyon Jung ·

    Self- and Other-Labels Induce Bidirectional Bias in LLM Judges

    arXiv:2608.18091v1 Announce Type: cross Abstract: As LLM-as-a-judge systems become increasingly widespread, self-preference in LLMs -- the tendency to favor one's own outputs -- raises growing concerns about evaluation reliability. However, it has been studied predominantly on ge…