PulseAugur
实时 14:58:02
English(EN) BadWAM: When World-Action Models Dream Right but Act Wrong

新的BadWAM框架揭示了世界动作模型中的漏洞

研究人员引入了BadWAM,这是一个用于评估世界动作模型(WAMs)对抗性攻击的框架。这些攻击利用微小的视觉扰动来破坏WAMs的预期未来与其执行动作之间的一致性。该框架包含两种类型的攻击:一种优先考虑任务失败,另一种则在诱导有害动作的同时保持一个看似合理的预期未来,这表明任务成功率显著下降。 AI

影响 凸显了具身AI系统潜在的安全漏洞,需要强大的防御机制。

排序理由 该集群包含一篇研究论文,详细介绍了评估特定类型AI模型对抗性攻击的新框架。

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

新的BadWAM框架揭示了世界动作模型中的漏洞

报道来源 [3]

  1. arXiv cs.LG TIER_1 English(EN) · Qi Li, Xingyi Yang, Xinchao Wang ·

    BadWAM:当世界-动作模型“做梦”正确但“行动”错误时

    arXiv:2607.15207v1 Announce Type: new Abstract: World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future world prediction. This coupling is of…

  2. arXiv cs.LG TIER_1 English(EN) · Xinchao Wang ·

    BadWAM:当世界-动作模型“做梦”是对的但“行动”是错的时候

    World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future world prediction. This coupling is often viewed as a source of robustness, interpreta…

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    BadWAM:当世界动作模型“做梦”对但“行动”错时

    World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future world prediction. This coupling is often viewed as a source of robustness, interpreta…