Researchers have developed a new adversarial attack method called Spatial Temporal Coherence Adversarial Attack (STCA) specifically designed to target black-box vision-language models (VLMs) used in autonomous driving systems. This attack operates in three stages: modality expansion for semantic frame selection, a spatial attack to create perturbations while maintaining similarity, and the STCA stage to disrupt temporal coherence using motion-guided masks. Experiments conducted on the BDD100K and nuScenes datasets using models like Video LLaVA-7B, Qwen2.5-VL-7B, and Dolphin demonstrated that current VLMs are highly vulnerable to such attacks, highlighting the urgent need for enhanced defenses. AI
IMPACT Highlights critical safety vulnerabilities in vision-language models used for autonomous driving, underscoring the need for robust defense mechanisms.
RANK_REASON The cluster contains a research paper detailing a novel adversarial attack method for vision-language models. [lever_c_demoted from research: ic=1 ai=1.0]
- autonomous driving
- BDD100K
- Dolphin
- Heyam Bin Jahlan
- nuScenes
- Qwen2.5-VL-7B
- Spatial Temporal Coherence Adversarial Attack
- vision-language model
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →