Researchers have identified a new jailbreak vulnerability in merged AI models, stemming from the underlying foundation model rather than the individual fine-tuned components. This "Basin-Aware Jailbreak" (BAJ) method generates adversarial suffixes that can bypass safety alignments across various merged models sharing the same backbone, even without knowledge of specific merging coefficients. Experiments demonstrate BAJ's effectiveness across different model families and its resilience against existing defenses. AI
IMPACT This research highlights a novel attack vector against merged AI models, potentially impacting the safety and reliability of systems built by combining multiple fine-tuned models.
RANK_REASON Research paper detailing a new AI security vulnerability. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →