PulseAugur
实时 11:01:31
English(EN) Broken Symmetry in LLM Refusal: Answer Release Is More Local Than Refusal Restoration

LLM拒绝机制呈现“对称性破坏”,答案仍可恢复

研究人员发现大型语言模型(LLM)在拒绝回答提示方面存在“对称性破坏”。他们发现,即使LLM生成了拒绝,通过局部干预仍然可以从其内部状态中恢复出正确答案。然而,重新施加拒绝是一个更复杂的过程,需要跨越多个位置进行更广泛的干预。这表明LLM的拒绝并非简单的开关,探针可恢复性可能会高估模型为安全和审计目的而控制行为的能力。 AI

影响 研究结果表明,当前的安全性探针可能会高估对LLM拒绝的控制能力,从而影响审计和对齐工作。

排序理由 学术论文发布在arXiv上,详细介绍了LLM行为的新发现。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

LLM拒绝机制呈现“对称性破坏”,答案仍可恢复

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Yiqi Liu, Yang Wang, Songxin Wang, Chenghao Xiao, Chenghua Lin ·

    LLM拒绝中的对称性破缺:回答释放比拒绝恢复更具局部性

    arXiv:2608.15772v1 Announce Type: new Abstract: When a language model refuses to answer a prompt, it is unclear whether the correct answer is erased from its internal representations, or merely suppressed at the output layer. We investigate this mechanism using a controlled withh…