Researchers have developed a novel method for detecting adversarial images by analyzing the response profiles of vision-language models (VLMs). This approach examines how a VLM interprets an image across a variety of general semantic prompts, summarizing these interactions using category-level statistics and prompt relationships. The detector, which keeps the VLM fixed, classifies these response profiles with a lightweight model, demonstrating strong performance against various known attacks and even those not encountered during training. The findings suggest that analyzing response patterns across semantic prompts offers a valuable supplementary signal for identifying adversarial inputs in frozen VLMs. AI
IMPACT This research offers a new technique for enhancing the security and reliability of vision-language models against malicious manipulation.
RANK_REASON Academic paper detailing a new method for detecting adversarial images using vision-language models. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- Influence Flower
- ScienceCast
- vision-language models
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →