PulseAugur
EN
LIVE 14:47:26

New research tackles LLM agent auditability and multi-agent safety risks

Two new research papers explore critical aspects of large language model (LLM) safety and enterprise application. The first paper introduces a "harness-engineering" approach to create auditable LLM agents with deterministic code, manifests, and validation artifacts, ensuring source-grounding and controlled behavior. The second paper proposes a controlled contrast design to disentangle safety risks in multi-agent LLM systems, differentiating between operational reframing, planner behavior, and delegation framing, and finding that reframing is a significant risk across models like GPT, Gemini, and DeepSeek, while Claude is more resistant. AI

IMPACT These papers offer new methodologies for improving the reliability and safety of LLM applications, particularly in enterprise and multi-agent settings.

RANK_REASON Two academic papers published on arXiv discussing LLM safety and engineering.

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 4 sources. How we write summaries →

New research tackles LLM agent auditability and multi-agent safety risks

COVERAGE [4]

  1. arXiv cs.AI TIER_1 English(EN) · Joongho Ahn, Moonsoo Kim ·

    From Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents

    arXiv:2607.08028v1 Announce Type: new Abstract: Enterprise large language model (LLM) applications often begin as prototypes whose behavior is carried by prompts and retrieval context. Productization adds requirements for source boundaries, entity routing, answer contracts, and r…

  2. arXiv cs.AI TIER_1 English(EN) · Lifei Liu, Haoran Yu, Xiaochong Jiang, Su Wang, Pin Qian, Yihang Chen ·

    Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety

    arXiv:2607.07097v1 Announce Type: new Abstract: Safety evaluations of multi-agent LLM systems often compare a direct prompt with a planner-executor pipeline and report the difference as a single "pipeline effect." We argue that this aggregate is difficult to interpret because it …

  3. arXiv cs.CL TIER_1 English(EN) · Moonsoo Kim ·

    From Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents

    Enterprise large language model (LLM) applications often begin as prototypes whose behavior is carried by prompts and retrieval context. Productization adds requirements for source boundaries, entity routing, answer contracts, and reproducible traces. We present a harness-enginee…

  4. arXiv cs.AI TIER_1 English(EN) · Yihang Chen ·

    Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety

    Safety evaluations of multi-agent LLM systems often compare a direct prompt with a planner-executor pipeline and report the difference as a single "pipeline effect." We argue that this aggregate is difficult to interpret because it conflates three mechanisms: harmful intent may b…