PulseAugur
中
实时 14:28:31
English(EN) Present but Rescaled: Chat-to-Agent Transfer of Additive Activation Steering

AI引导方法在代理部署中显示出不可预测的安全影响

一项新研究调查了加性激活引导从单轮聊天到ReAct代理的可迁移性,发现虽然引导方向一致地到达了后期层,但其行为影响是不可预测且依赖于模型的。研究表明,在某些模型上,代理部署可以将拒绝绕过向量放大高达2.00倍,而其他模型则显示出衰减,这表明安全不能被假定。这种分离表明重缩放发生在ReAct格式脚手架上,而不是工具观察上。 AI

影响 加性激活引导在代理AI部署中不可预测的安全结果,需要仔细的模型选择和安全评估。

排序理由 学术论文,详细介绍了一项关于AI模型行为和安全影响的新研究。

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

AI引导方法在代理部署中显示出不可预测的安全影响

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
学术论文,详细介绍了一项关于AI模型行为和安全影响的新研究。
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, safety, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
93 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [2]

  1. arXiv cs.LG TIER_1 English(EN) · Lucas Pinto ·

    存在但已缩放:加性激活引导的聊天到代理迁移

    arXiv:2607.09156v1 Announce Type: new Abstract: Additive activation steering (injecting a scaled residual-stream direction during generation) is calibrated almost entirely in single-turn chat, yet the models it targets are increasingly deployed as tool-using ReAct agents. We pres…

  2. arXiv cs.LG TIER_1 English(EN) · Lucas Pinto ·

    存在但已缩放:加性激活引导的聊天到代理迁移

    Additive activation steering (injecting a scaled residual-stream direction during generation) is calibrated almost entirely in single-turn chat, yet the models it targets are increasingly deployed as tool-using ReAct agents. We present the first systematic chat-to-agent transfer …