PulseAugur
实时 09:16:41
English(EN) Before You Poll with LLMs: A Deliberative Diagnostic Framework

新框架揭示大型语言模型无法准确模拟人类信念转变

一个名为审议式民意调查诊断框架(Deliberative Polling Diagnostic Framework)的新框架被引入,用于评估大型语言模型(LLMs)在接收新信息后如何更新其信念,这项能力对于它们模拟公众舆论至关重要。与以往的静态评估不同,该框架通过在相同的干预信息后比较人类和大型语言模型的信念转变来评估动态保真度。对五种前沿模型——GPT-5.1Gemini 2.0 FlashClaude Sonnet 4.5Llama 3.3-70BDeepSeek-V3 的测试表明,所有模型都未能准确模仿人类的审议过程,表现出信念逆转、过度反应或僵化等问题,这种现象被称为“自我谄媚”。 AI

影响 这项研究突显了大型语言模型在推理和信念更新方面的关键局限性,表明当前模型在模拟细微的公众舆论方面不可靠,可能需要显著改进其动态保真度。

排序理由 学术论文,介绍了一个用于大型语言模型的新评估框架。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新框架揭示大型语言模型无法准确模拟人类信念转变

本文如何被排名

Signal score
14 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
学术论文,介绍了一个用于大型语言模型的新评估框架。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, safety, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Ahmed Wali, Hassaan Tayyab ·

    在您使用大型语言模型进行民意调查之前:一个审议性诊断框架

    arXiv:2609.15849v1 Announce Type: cross Abstract: Can LLMs reason through new information like humans, or do they merely retrieve cached opinions? This is critical for silicon sampling, where LLM personas simulate public opinion at scale. Current evaluations test only whether per…