Researchers have developed a novel method for implanting programmable backdoors into Vision-Language Models (VLMs). Unlike previous static backdoor attacks, this new technique allows attackers to dynamically control target captions and generate corresponding triggers at inference time, even for unseen semantics. The attack involves a heuristic poisoning strategy to teach the model a general trigger-as-instruction rule, followed by a feature-space steganography method to map any target caption to a stealthy visual trigger. This approach maintains the model's clean utility and demonstrates effectiveness against common backdoor defenses. AI
IMPACT Introduces a new class of sophisticated attacks against VLMs, potentially impacting model security and deployment.
RANK_REASON Academic paper detailing a new method for implanting programmable backdoors in VLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →