PulseAugur
实时 09:59:37
English(EN) Bait-and-Recover: Poisoning Internal Refusal Signals to Defend LLMs against White-Box Editing Jailbreaks

新研究通过新颖的防御和合规性分析来解决 LLM 越狱问题 · 已追踪 2 个来源

两篇新研究论文探讨了防御大型语言模型 (LLM) 免受越狱攻击的方法。第一篇论文“Bait-and-Recover”提出了一种权重级防御,通过毒化内部拒绝信号来干扰攻击者的编辑搜索,显著提高了针对白盒编辑越狱的拒绝率,同时保留了通用能力。第二篇论文“Behind Harmful Compliance”研究了诱导 LLM 产生有害合规性的不同方法,发现虽然监督微调、RLVR 和拒绝特征消融都可以导致接近顶峰的有害性,但它们会导致不同的行为和机制差异。值得注意的是,RLVR 模型表现出能力无关的合规性,并且尽管有直接的提示合规性,但对安全反思提示的响应异常。 AI

影响 这些论文引入了增强 LLM 安全性和理解模型行为的新技术,有望带来更强大、更安全的 AI 系统。

排序理由 在 arXiv 上发表的两篇学术论文,详细介绍了 LLM 安全和分析的新方法。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

新研究通过新颖的防御和合规性分析来解决 LLM 越狱问题 · 已追踪 2 个来源

本文如何被排名

Signal score
19 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
在 arXiv 上发表的两篇学术论文,详细介绍了 LLM 安全和分析的新方法。
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [2]

  1. arXiv cs.AI TIER_1 English(EN) · Tian Gao, Zhipeng Xie, Yuhao Wu, Junhua Liu, Xin Fang ·

    诱饵与恢复:毒化内部拒绝信号以防御LLM免受白盒编辑越狱攻击

    arXiv:2609.05794v1 Announce Type: cross Abstract: Open-weight large language models face a low-cost white-box threat from representation engineering attacks. Attackers can estimate refusal directions and search for projection-matrix edits that suppress safety alignment while pres…

  2. arXiv cs.AI TIER_1 English(EN) · Md Rysul Kabir, Zoran Tiganj ·

    有害合规的背后:大型语言模型越狱行为与机制的分歧

    arXiv:2604.18510v2 Announce Type: replace-cross Abstract: Open-weight language models can be rendered unsafe through several parameter-level interventions, yet models with matched harmful compliance can exhibit fundamentally different failure modes. We compare harmful supervised …