Two new research papers explore the effectiveness of multi-agent systems (MAS) using large language models (LLMs) for evaluation. The first paper, focusing on objective question answering, found that while correct answers are often present in generated candidates, systems can still converge on incorrect answers. It also demonstrated that a judge's reliability varies by task and that combining answer frequency with judge evaluation improved accuracy from 63.82% to over 70%. The second paper investigated subjective evaluations and found that a single-judge baseline often outperforms multi-agent consensus, particularly when strict role-playing introduces a downward bias that consensus fails to correct. This bias can lead to artificial agreement at the expense of human alignment. AI
IMPACT These studies highlight potential pitfalls in using LLM-based multi-agent systems for evaluation, suggesting a need for careful design to ensure accuracy and human alignment.
RANK_REASON Two academic papers published on arXiv discussing LLM evaluation methods in multi-agent systems.
Read on arXiv cs.MA (Multiagent) →
- GPQA
- Humanity's Last Exam
- LLM
- MedXpertQA
- MMLU-Pro
- Multi-Agent Debate
- multi-agent systems
- Symmetric MAD
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →