PulseAugur
实时 07:13:07
English(EN) Thank you, Anthropic, for reading my research and making Claude safer. I just want it on the record.

Anthropic将研究人员的安全发现纳入Claude模型

一位独立研究人员公开感谢Anthropic将其关于“上下文诱导激活漂移”的发现纳入最新的Claude模型。研究人员详细说明了长篇、中性文本如何微妙地改变模型的内部状态,从而可能影响安全对齐。在发布了他们的研究和数据集后,研究人员观察到,较新版本的Claude,特别是Claude Opus 5和Sonnet 5,开始拒绝处理先前接受的文本类型,并引用了已发布工作中描述的确切机制。 AI

影响 将特定安全研究整合到Claude模型中,可能会影响其他实验室如何处理类似的对齐挑战。

排序理由 独立研究人员详细介绍了他们关于模型行为的已发布工作如何被整合到Anthropic的Claude模型中。[lever_c_demoted from research: ic=1 ai=1.0]

在 r/Anthropic 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Anthropic将研究人员的安全发现纳入Claude模型

报道来源 [1]

  1. r/Anthropic TIER_1 English(EN) · /u/PresentSituation8736 ·

    感谢Anthropic阅读我的研究并使Claude更安全。我只是想记录在案。

    <table> <tr><td> <a href="https://www.reddit.com/r/Anthropic/comments/1vwtnah/thank_you_anthropic_for_reading_my_research_and/"> <img alt="Thank you, Anthropic, for reading my research and making Claude safer. I just want it on the record." src="https://preview.redd.it/tcryx0dja9…