A new arXiv paper proposes a shift in how Large Language Models (LLMs) are used as judges, suggesting that verbalized confidence is now a more robust scoring mechanism than log-probabilities for post-2025 proprietary models. The research indicates that this "compatibility shift" is evident across various benchmarks like SummEval, AggreFact, and HelpSteer2, and involves up to 18 LLMs. The paper introduces an overconfidence advisory and self-debate to improve calibration and score distribution, noting that newer models accommodate these additions with minimal cost, unlike older models. AI
IMPACT Suggests a new standard for evaluating LLMs, potentially impacting how model capabilities are benchmarked and compared.
RANK_REASON Research paper published on arXiv detailing a new methodology for LLM evaluation. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →