PulseAugur
实时 09:05:04
English(EN) Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety

新研究应对LLM代理可审计性和多智能体安全风险

两篇新研究论文探讨了大语言模型(LLM)安全和企业应用的关键方面。第一篇论文介绍了一种“工具链工程”方法,通过确定性的代码、清单和验证工件来创建可审计的LLM代理,确保其来源可追溯和行为可控。第二篇论文提出了一种受控对比设计,以区分多智能体LLM系统中的安全风险,区分操作重构、规划器行为和委托框架,并发现重构是GPT、Gemini和DeepSeek等模型面临的重大风险,而Claude的抵抗力更强。 AI

影响 这些论文为提高LLM应用(尤其是在企业和多智能体环境中)的可靠性和安全性提供了新方法。

排序理由 两篇在arXiv上发表的学术论文,讨论LLM安全和工程。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 4 个来源。 我们如何撰写摘要 →

新研究应对LLM代理可审计性和多智能体安全风险

报道来源 [4]

  1. arXiv cs.AI TIER_1 English(EN) · Joongho Ahn, Moonsoo Kim ·

    从提示词到合同:利用工程技术打造可审计的企业级LLM代理

    arXiv:2607.08028v1 Announce Type: new Abstract: Enterprise large language model (LLM) applications often begin as prototypes whose behavior is carried by prompts and retrieval context. Productization adds requirements for source boundaries, entity routing, answer contracts, and r…

  2. arXiv cs.AI TIER_1 English(EN) · Lifei Liu, Haoran Yu, Xiaochong Jiang, Su Wang, Pin Qian, Yihang Chen ·

    多智能体LLM安全中的操作重构与批准框架委托

    arXiv:2607.07097v1 Announce Type: new Abstract: Safety evaluations of multi-agent LLM systems often compare a direct prompt with a planner-executor pipeline and report the difference as a single "pipeline effect." We argue that this aggregate is difficult to interpret because it …

  3. arXiv cs.CL TIER_1 English(EN) · Moonsoo Kim ·

    从提示词到合同:利用工程化技术打造可审计的企业级LLM代理

    Enterprise large language model (LLM) applications often begin as prototypes whose behavior is carried by prompts and retrieval context. Productization adds requirements for source boundaries, entity routing, answer contracts, and reproducible traces. We present a harness-enginee…

  4. arXiv cs.AI TIER_1 English(EN) · Yihang Chen ·

    多智能体LLM安全中的操作重构与批准框架委托

    Safety evaluations of multi-agent LLM systems often compare a direct prompt with a planner-executor pipeline and report the difference as a single "pipeline effect." We argue that this aggregate is difficult to interpret because it conflates three mechanisms: harmful intent may b…