This article details a new method for evaluating Large Language Models (LLMs) by introducing a third type of judge: the LLMs themselves, scoring each other's outputs. The author previously found that Gemini incorrectly flagged arithmetic as hallucination, prompting a test to see if this was a general LLM judge issue. By having Claude Sonnet 4.5 score GPT-5.5's outputs and vice versa, the evaluation now includes a "cross-judge" score. This new score is weighted alongside human and Gemini scores, with human judgment carrying the most weight (50%), to create a composite score that aims to provide a more reliable assessment of model performance, particularly regarding factual accuracy and reasoning. AI
IMPACT This cross-LLM evaluation method could lead to more robust and reliable benchmarks for assessing model capabilities, potentially influencing future model development and selection.
RANK_REASON The item describes a novel methodology for evaluating LLMs, including a new scoring approach and the use of LLMs as judges for each other. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →