PulseAugur
实时 09:44:38
English(EN) Moving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks

新的DRAA框架提升AI对抗白盒攻击的安全性

研究人员引入了一个名为动态路由自适应对齐(DRAA)的新框架,以增强大型基础模型(LFMs)在面对复杂的白盒攻击时的安全性。这些攻击针对内部安全机制,不同于之前的黑盒越狱。DRAA通过识别并动态绕过受损的安全路径来工作,确保模型即使在其主要安全路径中断时也能保持强大的拒绝行为和通用效用。实验表明,DRAA显著提高了模型对这些高级对抗技术的抵御能力。 AI

影响 增强AI模型对抗复杂对抗性攻击的鲁棒性,可能提高在实际部署中的安全性。

排序理由 该集群包含一篇详细介绍AI安全新技术的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的DRAA框架提升AI对抗白盒攻击的安全性

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Shangze Li, Chuancheng Shi, Simiao Xie, Lingzhi He, Cheng Ji, Zifeng Cheng, Fei Shen, Chao Wu, Tat-Seng Chua ·

    提升安全屏障:动态路由自适应对齐对抗白盒攻击

    arXiv:2608.02674v1 Announce Type: cross Abstract: With the widespread deployment of large foundation models (LFMs) in open environments, safety threats are shifting from black-box jailbreaks toward white-box attacks that directly identify and disrupt internal safety neurons or ro…