Researchers have developed HASTE, a novel multi-agent framework designed to automatically evolve agent harnesses against emerging cyber threats. This system addresses the challenge of adapting safety constraints when faced with limited evidence, such as brief descriptions or a few attack examples from threat reports. HASTE operates through an adversarial process where safety specifications are generated to guide harness updates, while attack cases are used to probe for remaining vulnerabilities. This iterative approach allows the framework to evolve harnesses effectively against new attacks, even those beyond the initially observed evidence, as demonstrated by consistent reductions in attack success rates across various models and attack types. AI
IMPACT This framework could enhance the robustness of AI systems against novel security threats by automating the adaptation of safety mechanisms.
RANK_REASON The cluster contains a research paper detailing a new framework for AI safety. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →