Researchers have developed Vis-Poison, a novel attack method that compromises multimodal retrieval-augmented generation (RAG) systems by poisoning the visual data itself. This attack bypasses traditional defenses by embedding malicious content directly into images, without altering associated text like captions or metadata. Vis-Poison has demonstrated significant success rates, ranging from 40.16% to 65.40% in black-box settings against large knowledge bases, and remains effective even against models capable of answering from their parametric knowledge alone. AI
IMPACT This research highlights a new vulnerability in multimodal AI systems, potentially impacting the security and reliability of AI applications that rely on visual data.
RANK_REASON The cluster describes a novel attack method detailed in an academic paper published on arXiv. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Hugging Face
- multimodal large language model
- multimodal retrieval-augmented generation
- Vis-Poison
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →