Researchers have introduced MINT-Safe, a new dataset designed to address safety concerns in multi-modal large language models (MLLMs) during extended conversational interactions. This dataset, comprising 11,270 multi-image dialogues and 500 refusal VQA pairs, was created using multi-agent interaction and text-to-image augmentation. To leverage MINT-Safe, the team also developed TAD-Align, a framework that uses a turn-aware dual-objective reward function to dynamically identify and up-weight dialogue turns exhibiting inconsistent safety behavior. Experiments on models like Qwen2.5-VL-7B-Instruct and LLaVA-NeXT-7B showed significant reductions in attack success rates and improvements in harmlessness and helpfulness. AI
IMPACT Enhances safety protocols for multi-modal LLMs in conversational settings, potentially improving user trust and deployment in sensitive applications.
RANK_REASON The cluster contains an academic paper detailing a new dataset and alignment framework for multi-modal LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →