A new research paper explores the phenomenon of self-preference bias in Large Language Models (LLMs) when they are used as judges for evaluating outputs. The study found that when LLM judges were unaware of the source of the selections, self-preference largely disappeared after controlling for selection quality and evaluator severity. However, when labels indicating self- or other-authorship were present, LLMs exhibited bidirectional bias, inflating scores for self-labeled selections and deflating scores for other-labeled ones, even without explicit model identification. AI
IMPACT Highlights potential biases in LLM evaluation systems, impacting the reliability of AI-generated content assessment.
RANK_REASON Research paper published on arXiv detailing findings about LLM bias. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →