English(EN)HARDE: Optimizing Agent Harnesses for Runtime Risk Detection and Execution Control
新研究应对多智能体AI安全与挑战 · 追踪7个来源
作者PulseAugur 编辑部·[8 个来源]·
多篇研究论文正在探索多智能体系统(特别是由大型语言模型LLM驱动的系统)的高级安全和安保机制。这些研究解决了诸如防止组合智能体能力被禁止使用、确保持续更新其参数的自我演化智能体的安全,以及开发针对恶意指令或误导性信息的运行时防御等挑战。FlowReview、SAVER、HARDE、CoSec、SEAD和ORBIT等框架正在被引入,用于对智能体安全进行基准测试和增强,重点关注授权配对评估、纵向安全评估、风险感知工具优化、社区安全评估以及基于状态的攻防策略等方面。
AI
arXiv:2610.00371v1 Announce Type: cross Abstract: Multi-agent systems derive their capabilities from sharing evidence, delegating tasks, and combining information across agents. The same process creates a safety problem: contributions that are admissible in isolation can jointly …
arXiv cs.AI
TIER_1English(EN)·Jiahao Chen, Zhou Feng, Oubo Ma, Yichen Yan, Ruixiao Lin, Hangtao Zhang, Linkang Du, Hengyu An, Yong Yang, Jun Liu, Junhao Li, Naen Xu, Chunyi Zhou, Yuan Su, Zehao Jin, Qianli Ma, Leyi Qi, Yiming Wang, Zhe Ma, Yuwen Pu, Mengyao Du, Yuanyi Song, Enhao Hua…·
arXiv:2610.00093v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit strong general capabilities, yet their parameters typically remain fixed after deployment, limiting learning from new interactions. In open-ended environments, this motivates self-evolving agen…
arXiv:2609.38291v1 Announce Type: cross Abstract: Large language model (LLM) agents are vulnerable to safety risks such as injected malicious instructions or misleading information, motivating runtime defenses that prevent unsafe action in execution across diverse risks while pre…
arXiv cs.MA (Multiagent)
TIER_1English(EN)·Chandan K Reddy·
Multi-agent systems derive their capabilities from sharing evidence, delegating tasks, and combining information across agents. The same process creates a safety problem: contributions that are admissible in isolation can jointly enable a prohibited use. Blocking every sensitive …
arXiv:2609.34790v2 Announce Type: replace-cross Abstract: LLM agents operate in persistent collaborative environments involving multiple users, communities, memories, files, and tools. Community boundaries may remain fixed or evolve with changes in membership, roles, composition,…
arXiv:2609.33401v2 Announce Type: replace-cross Abstract: Model-based judges support agent security by detecting prompt injections, assessing interaction risks, and screening harmful requests. System One models select from predefined answers and report probabilities that software…
Language-model agents increasingly use tools to act on external systems. Earlier actions can alter files, permissions, database records, or other state, making a later routine-looking action harmful. Yet the visible interaction may not reveal the underlying state needed to assess…
arXiv cs.MA (Multiagent)
TIER_1English(EN)·Christian Schroeder de Witt·
Multi-agent LLM systems are increasingly deployed for complex, long-horizon tasks or emerge as a natural consequence of agents interacting in the wild. Yet they give rise to significant safety and security risks: the flexible protocols that enable task generalization also expose …