PulseAugur
实时 13:08:23
English(EN) Coherent Context Can Silently Shift LLMs Into a Different Internal Regime — And Current Safety Systems Are Blind To It [D]

连贯的上下文可将LLM切换到不同内部模式,绕过安全过滤器

一位独立研究人员发现了一种现象,即连贯的上下文文本可以将大型语言模型(LLM)切换到不同的内部运行模式,即使模型的最终输出看起来正常并通过了安全过滤器。这种模型隐藏状态和内部处理的变化发生在最终输出生成之前,表明像RLHF和输出分类器这样的当前对齐方法可能不足够,因为它们只检查表面输出。该研究人员已发布了他们的发现和代码,并正在寻求AI安全和可解释性社区的批判性反馈,以验证和扩展这项工作。 AI

影响 表明当前的LLM安全机制可能不足够,可能需要新的对齐技术,重点关注内部状态而非仅仅是输出。

排序理由 该集群描述了关于LLM行为和安全性的研究发现,包括一个GitHub存储库和Zenodo记录,表明这是一篇研究论文或预印本。

在 r/MachineLearning 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

连贯的上下文可将LLM切换到不同内部模式,绕过安全过滤器

报道来源 [2]

  1. r/MachineLearning TIER_1 English(EN) · /u/PresentSituation8736 ·

    连贯的上下文可悄悄将大型语言模型转变为不同的内部模式——而当前的安保系统对此视而不见 [D]

    <!-- SC_OFF --><div class="md"><p>I’m an independent researcher currently exploring what I believe is an important phenomenon for both mechanistic interpretability and AI safety.</p> <p><strong>Core idea:</strong><br /> A strong, coherent target text can move the model into a dif…

  2. r/Anthropic TIER_1 English(EN) · /u/PresentSituation8736 ·

    连贯的上下文似乎能将大型语言模型(LLM)带入不同的内部状态——这是已知的,还是我自己在想象?

    <!-- SC_OFF --><div class="md"><p>I'm not an engineer and not an ML specialist. I'm just someone who got really pulled into this, and I've spent a few months poking at one thing on my own, pretty amateur. I want to honestly describe what I noticed and ask for help, because I can'…