Researchers have introduced MEDLEY-BENCH, a new benchmark designed to evaluate how large language models revise their beliefs when presented with conflicting evidence or social pressure. Unlike benchmarks that focus solely on final answer quality, MEDLEY-BENCH assesses structured private self-review and analyst-conditioned social revision. Initial evaluations of 35 models across 12 families revealed varied performance in belief revision, with a notable pattern of lower scores in the 'Evaluation-mapped' composite across most models. AI
IMPACT This benchmark could lead to more robust LLM evaluations that consider belief revision and social influence.
RANK_REASON The cluster contains a research paper introducing a new benchmark for evaluating LLM behavior. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →