Two new research papers explore vulnerabilities in AI models, focusing on jailbreak attacks. The first paper examines input-side defenses against semantic attacks on locally deployed Large Language Models (LLMs), identifying specific assumptions that defenses rely on and testing their failure points across various open-weight models. The second paper introduces a framework called Text-Anchored Semantic Perturbation Attack (TA-SPA) designed to exploit Multimodal Large Language Models (MLLMs) by optimizing transferable perturbations in a text-anchored semantic space, demonstrating effectiveness against commercial MLLMs. AI
IMPACT Highlights ongoing challenges in AI safety and alignment, particularly for locally deployed and multimodal models, necessitating further research into robust defense mechanisms.
RANK_REASON Two academic papers published on arXiv detailing new methods for jailbreaking AI models and analyzing existing defenses.
- arXiv
- Large Language Models
- Multimodal Large Language Models
- TA-SPA
- Text-Anchored Semantic Perturbation Attack
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →