Researchers have identified a new safety vulnerability in multimodal large language models (MLLMs) used for video moderation, termed Distributed Implicit Harm (DIH). This occurs when seemingly harmless video components combine to create an overall harmful message, a phenomenon that current MLLMs struggle to detect. The study introduces a framework to generate over 9,000 DIH videos with annotations and benchmarks over 30 MLLMs, revealing significant detection deficits even in frontier models. AI
IMPACT Highlights a critical safety blind spot in AI video moderation, potentially impacting content safety systems and requiring new detection methods.
RANK_REASON Academic paper detailing a new safety vulnerability in AI models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →