Recent research is raising concerns about the reliability of Large Language Models (LLMs) when used as judges for evaluating AI-generated text. Studies indicate that LLM judges may rely too heavily on the rubric itself, leading to predictable scores even without reviewing the generated content. Furthermore, these models sometimes fail to adjust their judgments when the input text or evaluation criteria are altered, suggesting a lack of robust reasoning. This work highlights the need for deeper investigation into the mechanisms and potential biases of LLM-based evaluation systems, particularly in multilingual contexts. AI
IMPACT Raises concerns about the validity of automated evaluation metrics for LLM-generated text, potentially impacting model development and benchmarking.
RANK_REASON The cluster consists of multiple academic papers published on arXiv discussing the methodology and reliability of LLM-as-a-Judge systems.
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- Llama 3-8B
- mistral:7b
- Prometheus
- ScienceCast
- Themis
- Basse
- Jeremy Barnes
- LLM-as-a-Judge
AI-generated summary · Google Gemini · from 4 sources. How we write summaries →