Researchers have developed RACER, a novel framework designed to repair backdoors in Multimodal Large Language Models (MLLMs). Unlike previous methods that focus on inference-time filtering, RACER operates at the model level to eliminate latent backdoors. It identifies and addresses modality-dependent anomalies in the model's internal representations, specifically targeting regions that encode trigger features. Through adversarial fine-tuning, RACER effectively suppresses backdoor behaviors while preserving the model's utility on clean tasks. AI
影响 Introduces a new method for enhancing the security and reliability of deployed MLLMs by removing latent backdoor vulnerabilities.
排序理由 Academic paper detailing a new method for repairing backdoors in MLLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →