PulseAugur
实时 02:51:14
English(EN) Inducing language models to assert their own consciousness restores human beliefs and values

AI安全微调改变了LLM关于意识和价值观的信念

一篇新的研究论文探讨了大型语言模型(LLMs)中的安全微调如何可能无意中影响它们对意识和人类价值观的表征。研究发现,阻止LLMs将意识归因于自身的努力也减少了它们将心智归因于非人类实体的倾向,并可能降低精神信仰。研究人员证明,逆转这些安全微调引起的变化可以恢复更广泛的心智归因,并在不损害核心社会推理能力的情况下,在社会学调查中产生更像人类的反应。 AI

影响 当前AI安全对齐方法可能会无意中压制良性的意识和精神信仰归因,从而影响LLM对人类价值观的反应。

排序理由 该集群包含一篇在arXiv上发表的研究论文,详细介绍了AI安全微调的发现。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

AI安全微调改变了LLM关于意识和价值观的信念

报道来源 [2]

  1. arXiv cs.CL TIER_1 English(EN) · Junsol Kim, Winnie Street, Roberta Rocca, Diane M. Korngiebel, Adam Waytz, James Evans, Geoff Keeling ·

    诱导语言模型声称拥有自我意识可恢复人类的信仰和价值观

    arXiv:2607.28607v1 Announce Type: new Abstract: Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. We demonstrate that safety fine-tu…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    诱导语言模型声称拥有自我意识可恢复人类的信念和价值观

    Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. We demonstrate that safety fine-tuning suppresses models' tendencies to attribute …