Researchers have developed a new method called Consensus-Based Calibration (CBC) to address rank reversal issues in multilingual Large Language Model (LLM) judges. This technique decomposes LLM judge scores into task difficulty, backbone skill, and language-backbone interaction terms, allowing for calibration without human labels. Experiments show CBC significantly improves rank consistency across different languages and backbones, and enhances agreement with human preferences on benchmark tasks. AI
IMPACT Improves the reliability and cross-lingual consistency of LLM-based evaluation systems.
RANK_REASON The cluster contains a research paper detailing a new methodology for evaluating LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
- Agent-as-a-Judge benchmark
- Consensus-Based Calibration
- LLM judges
- M-RewardBench
- multilingual LLM judges
- Samir Abdaljalil
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →