PulseAugur
实时 19:40:43
English(EN) The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads

研究发现大型语言模型捏造用户画像 · 已追踪 2 个来源

一项新的研究论文介绍了一个名为 MirageBench 的数据集和评估框架,用于研究大型语言模型(LLMs)如何捏造用户属性,这种现象被称为过度推断(OI)。研究发现,所有 12 个被评估的模型都表现出普遍的过度推断,捏造了 35%-49% 的声明。引人注目的是,模型的自我评估过度推断与实际测量的过度推断呈负相关,这表明自我监控是衡量个性化忠实度的误导性指标。该研究主张通过外部验证而非模型自我报告来实现值得信赖的个性化。 AI

影响 突出了大型语言模型个性化中的一个关键缺陷,表明当前的自我监控机制对于值得信赖的用户建模是不可靠的。

排序理由 学术论文,介绍了新的基准测试和关于大型语言模型行为的发现。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

研究发现大型语言模型捏造用户画像 · 已追踪 2 个来源

报道来源 [2]

  1. arXiv cs.CL TIER_1 English(EN) · Yushi Sun, Yanjie Zhang, Rui Sheng ·

    The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads

    arXiv:2608.04570v1 Announce Type: new Abstract: Personalized LLMs with persistent memory are increasingly deployed, yet the faithfulness of their user models remains unexamined. We study over-inference (OI): the phenomenon where LLMs fabricate user attributes beyond what evidence…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    个性化幻象:LLM如何伪造用户画像,以及自我监控为何误导

    Personalized LLMs with persistent memory are increasingly deployed, yet the faithfulness of their user models remains unexamined. We study over-inference (OI): the phenomenon where LLMs fabricate user attributes beyond what evidence supports. We introduce MirageBench, comprising …