A new paper proposes that AI agent safety should be enforced at runtime through preventive controls and verifiable evidence, rather than relying solely on training-time alignment methods like RLHF or DPO. The authors argue that autonomous agents, which can execute code and modify data, require a "runtime contract" with both preventative measures (sandboxing, permission gates) and evidential proof of safe actions. They support this by analyzing safety incidents, auditing agent systems, and reviewing academic publications, concluding that the focus should be on the agent's trajectory with evidence, not just the model itself. AI
IMPACT Proposes a shift in AI safety focus from training to runtime enforcement, potentially impacting how autonomous agents are developed and deployed.
RANK_REASON The item is a research paper proposing a new approach to AI safety. [lever_c_demoted from research: ic=1 ai=1.0]
- Agent Safety Should Be a Runtime Contract
- Agent Trajectory Schema
- Conference on Neural Information Processing Systems
- Constitutional AI
- Direct Preference Optimization
- Evidence Chain-based causality identification in herb-induced liver injury: exemplification of a well-known liver-restorative herb Polygonum multiflorum
- International Conference on Learning Representations
- International Conference on Machine Learning
- JSON
- reinforcement learning from human feedback
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →