A new paper argues that AI safety should be enforced through runtime contracts rather than solely during the training phase. The authors propose a two-pronged approach: a preventive face that blocks dangerous actions before they occur using sandboxes and filters, and an evidential face that requires verifiable proof of safe actions. This perspective is supported by evidence from AI safety incidents, audits of agent systems, and a review of academic publications, suggesting that agentic AI faces similar pressures to communities like computer security and experimental science, which have adopted runtime contracts. AI
IMPACT This research suggests a shift in AI safety focus from model training to runtime enforcement, potentially impacting how AI agents are developed and deployed.
RANK_REASON The cluster contains an academic paper proposing a new approach to AI safety.
- Agent Safety Should Be a Runtime Contract
- Agent Trajectory Schema
- Conference on Neural Information Processing Systems
- Constitutional AI
- Direct Preference Optimization
- Evidence chain-based causality identification in herb-induced liver injury: exemplification of a well-known liver-restorative herb Polygonum multiflorum
- International Conference on Learning Representations
- International Conference on Machine Learning
- JSON
- reinforcement learning from human feedback
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →