Researchers are developing new methods to address scoring bias in Large Language Models (LLMs) when they are used as judges for evaluating text quality. One approach involves instructing LLMs to generate random numbers to identify and rectify latent numerical biases, showing improved performance over existing methods. Another study introduces JudgeArena, a unified framework that standardizes LLM-judge evaluations across various benchmarks and models, enhancing reproducibility and transparency. Additionally, a new benchmark called JudgeBiasBench systematically quantifies different types of biases in LLM judges, and a technique called Chain-of-Models uses a secondary LLM to audit the reasoning of a primary LLM judge, demonstrating improved robustness against specific cognitive biases. AI
IMPACT These advancements aim to improve the reliability and reproducibility of LLM-based evaluations, crucial for model development and comparison.
RANK_REASON Multiple research papers proposing new methods and frameworks for evaluating LLM judges.
Read on Hugging Face Daily Papers →
- arXiv
- Chain-of-Models
- GLM-5
- GPT-4o
- Hugging Face
- Kimi K2.5
- Qwen2.5-72B
- alphaXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Hongli Zhou
- JudgeBiasBench
- ScienceCast
- AlpacaEval
- Arena-Hard
- JudgeArena
- LLM-as-a-Judge
- m-Arena-Hard
AI-generated summary · Google Gemini · from 7 sources. How we write summaries →