PulseAugur
EN
LIVE 08:18:11

New benchmark MEDLEY-BENCH tests LLM belief revision under social pressure

Researchers have introduced MEDLEY-BENCH, a new benchmark designed to evaluate how large language models revise their beliefs when presented with conflicting evidence or social pressure. Unlike benchmarks that focus solely on final answer quality, MEDLEY-BENCH assesses structured private self-review and analyst-conditioned social revision. Initial evaluations of 35 models across 12 families revealed varied performance in belief revision, with a notable pattern of lower scores in the 'Evaluation-mapped' composite across most models. AI

IMPACT This benchmark could lead to more robust LLM evaluations that consider belief revision and social influence.

RANK_REASON The cluster contains a research paper introducing a new benchmark for evaluating LLM behavior. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New benchmark MEDLEY-BENCH tests LLM belief revision under social pressure

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Farhad Abtahi, Abdolamir Karbalaie, Eduardo Illueca-Fernandez, Fernando Seoane ·

    MEDLEY-BENCH: Benchmarking Behavioural Metacognition and Belief Revision Under Social Pressure in Large Language Models

    arXiv:2604.16009v2 Announce Type: replace Abstract: Most large language model benchmarks evaluate final-answer quality but reveal little about how models revise beliefs under disagreement or conflicting evidence. We introduce MEDLEY-BENCH, an open benchmark comparing structured p…