PulseAugur
EN
LIVE 12:31:20

New AI defenses and attacks target vision-language models

Researchers have developed new methods to defend against and exploit backdoor attacks in advanced AI models. One approach, BYORn, aims to improve the robustness of large vision-language models by identifying and replacing semantically implausible responses generated during fine-tuning, thereby preserving clean task performance. In parallel, a new attack called SILENTDRIFT targets vision-language-action models used in robotics, exploiting action chunking to create stealthy backdoor attacks that are visually indistinguishable from legitimate demonstrations. AI

IMPACT Highlights ongoing research into AI security, with implications for model robustness and the potential for adversarial manipulation in critical applications.

RANK_REASON Two research papers detailing new methods for defending against and exploiting backdoor attacks in AI models.

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

New AI defenses and attacks target vision-language models

COVERAGE [2]

  1. arXiv cs.LG TIER_1 English(EN) · Ivan Saboli\'c, Marin Or\v{s}i\'c, Josip \v{S}ari\'c, Sven Lon\v{c}ari\'c ·

    BYORn: Bootstrap Your Own Responses to Defend Large Vision-Language Models Against Backdoor Attacks

    arXiv:2606.02947v1 Announce Type: new Abstract: Supervised fine-tuning is the predominant approach for adapting autoregressive vision-language models to downstream tasks. Recent work has shown that this paradigm is highly vulnerable to backdoor attacks, and that existing defenses…

  2. arXiv cs.AI TIER_1 English(EN) · Bingxin Xu, Yuzhang Shang, Binghui Wang, Emilio Ferrara ·

    SilentDrift: Exploiting Action Chunking for Stealthy Backdoor Attacks on Vision-Language-Action Models

    arXiv:2601.14323v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models are increasingly deployed in safety-critical robotic applications, yet their security vulnerabilities remain underexplored. We identify a fundamental security flaw in modern VLA systems:…