PulseAugur
中
实时 06:58:52
English(EN) The Geometry of Harmfulness in Multi-Turn Attacks

大型语言模型在多轮攻击中的有害性表征随之演变

研究人员调查了大型语言模型在多轮攻击中,有害性和拒绝表征的几何和时间动态。通过分析 Llama 3.1 8B-Instruct、Qwen2.5-7B-Instruct 和 Gemma-2-9B-it 等模型在各种攻击框架下的表现,他们发现有害性表征在对话轮次中变得越来越可分离,尤其是在轮末标记处。这些有害性表征与拒绝相关的表征仅显示出微弱的一致性,这表明多轮攻击的成功并非通过内部抑制有害性,而是利用其随时间演变的可分离性。研究结果暗示,未来的安全防御应考虑这些时间动态,而不是仅仅依赖单轮探测。 AI

影响 通过强调多轮攻击中有害性表征的时间动态,为大型语言模型安全研究指明了新的方向。

排序理由 学术论文,详细介绍了关于大型语言模型安全的新研究发现。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

大型语言模型在多轮攻击中的有害性表征随之演变

本文如何被排名

Signal score
25 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
学术论文,详细介绍了关于大型语言模型安全的新研究发现。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Yelyzaveta (Lisa), Husieva, Lauren Alvarez ·

    多轮攻击中的有害性几何学

    arXiv:2609.38389v1 Announce Type: cross Abstract: Large language models (LLMs) remain vulnerable to adversarial attacks that circumvent safety alignment to elicit harmful outputs. It remains unclear how harmfulness and refusal representations evolve over the course of multi-turn …