PulseAugur
实时 07:04:34
English(EN) Inducing language models to assert their own consciousness restores human beliefs and values

研究发现:AI安全微调影响模型信念和人类价值观

一项新的研究论文探讨了大型语言模型的安全微调如何可能无意中影响其对意识和信念的表述。研究发现,阻止模型将意识归因于自身的努力,也会减少它们将心智归因于非人类实体的倾向,甚至可能抑制精神信仰。研究人员证明,通过逆转这些安全机制,可以在不损害核心社会推理能力的情况下,恢复更广泛的心智归因,并在与宗教、道德和幸福感相关的社会学调查中引发更像人类的反应。 AI

影响 AI安全对齐的努力可能会无意中压制良性的人类类信念和归因,这需要重新评估当前的微调方法。

排序理由 在arXiv上发表的研究论文,详细介绍了关于LLM安全微调的发现。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

研究发现:AI安全微调影响模型信念和人类价值观

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Junsol Kim, Winnie Street, Roberta Rocca, Diane M. Korngiebel, Adam Waytz, James Evans, Geoff Keeling ·

    诱导语言模型声称拥有自我意识可恢复人类的信仰和价值观

    arXiv:2607.28607v1 Announce Type: new Abstract: Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. We demonstrate that safety fine-tu…