PulseAugur
中
实时 09:11:24
English(EN) HARDE: Optimizing Agent Harnesses for Runtime Risk Detection and Execution Control

新研究应对多智能体AI安全与挑战 · 追踪7个来源

多篇研究论文正在探索多智能体系统(特别是由大型语言模型LLM驱动的系统)的高级安全和安保机制。这些研究解决了诸如防止组合智能体能力被禁止使用、确保持续更新其参数的自我演化智能体的安全,以及开发针对恶意指令或误导性信息的运行时防御等挑战。FlowReview、SAVER、HARDE、CoSec、SEAD和ORBIT等框架正在被引入,用于对智能体安全进行基准测试和增强,重点关注授权配对评估、纵向安全评估、风险感知工具优化、社区安全评估以及基于状态的攻防策略等方面。 AI

影响 这些研究工作旨在为日益复杂的AI智能体系统建立强大的安全和安保协议,这对于它们在现实世界应用中的可靠部署至关重要。

排序理由 多篇在arXiv上发表的学术论文,详细介绍了AI智能体安全与安保的新框架和基准。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 8 个来源。 我们如何撰写摘要 →

新研究应对多智能体AI安全与挑战 · 追踪7个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
多篇在arXiv上发表的学术论文,详细介绍了AI智能体安全与安保的新框架和基准。
Source corroboration
8 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
7 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [8]

  1. arXiv cs.AI TIER_1 English(EN) · Yunbei Zhang, Saiyue Lyu, Janet Wang, Yingqiang Ge, Jiang Guo, Jihun Hamm, Chandan K Reddy ·

    不禁用即可拒绝:多智能体系统的授权配对评估与控制

    arXiv:2610.00371v1 Announce Type: cross Abstract: Multi-agent systems derive their capabilities from sharing evidence, delegating tasks, and combining information across agents. The same process creates a safety problem: contributions that are admissible in isolation can jointly …

  2. arXiv cs.AI TIER_1 English(EN) · Jiahao Chen, Zhou Feng, Oubo Ma, Yichen Yan, Ruixiao Lin, Hangtao Zhang, Linkang Du, Hengyu An, Yong Yang, Jun Liu, Junhao Li, Naen Xu, Chunyi Zhou, Yuan Su, Zehao Jin, Qianli Ma, Leyi Qi, Yiming Wang, Zhe Ma, Yuwen Pu, Mengyao Du, Yuanyi Song, Enhao Hua… ·

    自主演化Agent中的安全性:一项调查

    arXiv:2610.00093v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit strong general capabilities, yet their parameters typically remain fixed after deployment, limiting learning from new interactions. In open-ended environments, this motivates self-evolving agen…

  3. arXiv cs.CL TIER_1 English(EN) · Zhuo Liu, Moxin Li, Zhixin Ma, Wentao Shi, Wenjie Wang, Fuli Feng ·

    HARDE:优化代理工具以实现运行时风险检测和执行控制

    arXiv:2609.38291v1 Announce Type: cross Abstract: Large language model (LLM) agents are vulnerable to safety risks such as injected malicious instructions or misleading information, motivating runtime defenses that prevent unsafe action in execution across diverse risks while pre…

  4. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Chandan K Reddy ·

    拒绝而不禁用:多智能体系统的授权配对评估与控制

    Multi-agent systems derive their capabilities from sharing evidence, delegating tasks, and combining information across agents. The same process creates a safety problem: contributions that are admissible in isolation can jointly enable a prohibited use. Blocking every sensitive …

  5. arXiv cs.AI TIER_1 English(EN) · Hao Chen, Wenhui Dong, Ye Chen, Jiezhi Yao, Chenbo Xia, Yuwen Qu, Renxiang Wang, Fudong Yuan, Camil Hamami, Chenglong Pan, Xinquan Yue, Ziyu Wang, Fengyu Ye, Chenyang Si, Caifeng Shan ·

    CoSec:社区中代理安全性的基准测试

    arXiv:2609.34790v2 Announce Type: replace-cross Abstract: LLM agents operate in persistent collaborative environments involving multiple users, communities, memories, files, and tools. Community boundaries may remain fixed or evolve with changes in membership, roles, composition,…

  6. arXiv cs.AI TIER_1 English(EN) · Yixuan Liu ·

    评估 System One 模型在代理安全决策中的可靠性、校准和选择性自动化

    arXiv:2609.33401v2 Announce Type: replace-cross Abstract: Model-based judges support agent security by detecting prompt injections, assessing interaction risks, and screening harmful requests. System One models select from predefined answers and report probabilities that software…

  7. Hugging Face Daily Papers TIER_1 English(EN) ·

    SEAD:一种基于状态的工具使用Agent攻防视角

    Language-model agents increasingly use tools to act on external systems. Earlier actions can alter files, permissions, database records, or other state, making a later routine-looking action harmful. Yet the visible interaction may not reveal the underlying state needed to assess…

  8. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Christian Schroeder de Witt ·

    ORBIT:一个用于多智能体安全与安保评估的框架

    Multi-agent LLM systems are increasingly deployed for complex, long-horizon tasks or emerge as a natural consequence of agents interacting in the wild. Yet they give rise to significant safety and security risks: the flexible protocols that enable task generalization also expose …