A new framework called the Deliberative Polling Diagnostic Framework has been introduced to evaluate how Large Language Models (LLMs) update their beliefs in response to new information, a capability crucial for their use in simulating public opinion. Unlike previous static evaluations, this framework assesses dynamic fidelity by comparing human and LLM belief shifts after identical informational interventions. Testing five frontier models—GPT-5.1, Gemini 2.0 Flash, Claude Sonnet 4.5, Llama 3.3-70B, and DeepSeek-V3—revealed that all models failed to accurately mimic human deliberation, exhibiting issues such as belief reversal, overshoot, or rigidity, a phenomenon termed self-sycophancy. AI
IMPACT This research highlights critical limitations in LLM reasoning and belief updating, suggesting current models are unreliable for simulating nuanced public opinion and may require significant improvements in dynamic fidelity.
RANK_REASON Academic paper introducing a new evaluation framework for LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
- America in One Room
- Claude Sonnet 4.5
- DeepSeek-V3
- Deliberative Polling Diagnostic Framework
- Gemini 2.0 Flash
- GPT-5.1
- Llama 3.3-70B
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →