PulseAugur
EN
LIVE 08:54:09

New PoisonVID attack bypasses safety features in VideoLLMs

Researchers have developed a novel attack called PoisonVID that can bypass safety measures in Video Large Language Models (VideoLLMs). These models are used to moderate user-generated content by sampling key frames, but PoisonVID can manipulate this sampling process. The attack works by subtly altering harmful video frames so that the VideoLLM's prompt-guided sampling mechanism fails to identify them, effectively suppressing safety alerts. This method has demonstrated high success rates across various VideoLLM architectures and harmful content categories, even surviving several defense mechanisms. AI

IMPACT This research highlights a critical vulnerability in current VideoLLM safety systems, potentially impacting content moderation and requiring new defense strategies.

RANK_REASON Academic paper detailing a new attack method against AI models. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CV →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New PoisonVID attack bypasses safety features in VideoLLMs

COVERAGE [1]

  1. arXiv cs.CV TIER_1 English(EN) · Yuxin Cao, Wei Song, Jingling Xue, Jin Song Dong ·

    Poisoning Prompt-Guided Sampling in Video Large Language Models

    arXiv:2509.20851v2 Announce Type: replace Abstract: Video Large Language Models (VideoLLMs) are increasingly deployed as automated moderators on user-generated video platforms, where a few unwatched seconds of harmful footage are enough to suppress a safety alert. Because encodin…