JudgeBench: A Benchmark for Evaluating LLM-based Judges
PulseAugur coverage of JudgeBench: A Benchmark for Evaluating LLM-based Judges — every cluster mentioning JudgeBench: A Benchmark for Evaluating LLM-based Judges across labs, papers, and developer communities, ranked by signal.
-
LLM-as-a-Judge models show significant reliability and bias issues, study finds
A new study evaluating LLM-as-a-Judge models reveals significant issues with their reliability and validity. The research, which analyzed 21 judges across multiple benchmarks and over 541,000 judgments, found that commo…
-
AdaJudge framework improves LLM reward modeling with adaptive pooling
Researchers have introduced AdaJudge, a novel framework designed to enhance the accuracy of reward modeling in large language models. This approach tackles limitations in current static pooling strategies by adapting bo…
-
RTLC prompting boosts LLM judge accuracy by 14 percentage points
Researchers have developed a new three-stage prompting technique called RTLC (Research, Teach-to-Learn, Critique) that significantly improves the accuracy of large language models when used as judges. This method, inspi…