English(EN)Fully Unleashing the Multimodal Attacker: Meta-Adaptive Jailbreaking of Vision-Language Models
新研究探讨了大型语言模型和视觉-语言模型的先进越狱技术及检测方法
作者PulseAugur 编辑部·[10 个来源]·
研究人员正在开发先进的方法来测试大型语言模型和视觉-语言模型在面对越狱尝试时的安全性和鲁棒性。SEAV等新框架专注于验证模型响应的正确性和程序准确性,而不仅仅是语义上的合理性。其他研究引入了元自适应攻击,该攻击会优化攻击者本身以利用多模态模型的漏洞,在GPT-4o和Gemini 3 Pro Preview等领先模型上取得了高成功率。此外,还在探索通过分析模型激活来检测未见过越狱攻击的方法,旨在提高安全评估的泛化能力和效率。
AI
arXiv:2609.03781v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used in multilingual settings, yet their safety is still evaluated primarily in English. This limits our understanding of how alignment failures manifest in low-resource and culturally…
Large language models (LLMs) are increasingly used in multilingual settings, yet their safety is still evaluated primarily in English. This limits our understanding of how alignment failures manifest in low-resource and culturally diverse languages. We introduce IndicSafeEval, a …
arXiv:2609.00498v1 Announce Type: new Abstract: Jailbreak robustness has become central to large language model (LLM) safety evaluation, yet prevailing methodologies rely primarily on refusal behavior, semantic resemblance, and intent-matching heuristics that emphasize linguistic…
The safety of large vision-language models is increasingly stress-tested by multimodal jailbreaks, yet existing attacks remain largely static at the meta level: template-based attacks freeze the image--text layout, while iterative attacks adapt only the image--text content with f…
Multimodal Large Language Models (MLLMs) have achieved remarkable progress in vision-language interaction, yet their safety alignment remains vulnerable to jailbreak attacks. A key challenge is that safety behavior learned in the textual space does not reliably transfer to fused …
Large vision-language models (LVLMs) have achieved remarkable progress in video understanding and reasoning. Despite extensive studies on text- and image-based jailbreaks, video jailbreaks against LVLMs remain largely unexplored. Existing video jailbreak methods mainly manipulate…
arXiv:2608.27531v1 Announce Type: cross Abstract: The safety of large vision-language models is increasingly stress-tested by multimodal jailbreaks, yet existing attacks remain largely static at the meta level: template-based attacks freeze the image--text layout, while iterative…
arXiv:2608.26971v1 Announce Type: new Abstract: In recent years, image-to-video (I2V) generation models have made remarkable progress in subject consistency and temporal coherence, enabling high quality video synthesis. However, these advances also introduce new safety risks. Exi…
arXiv:2508.09201v5 Announce Type: replace-cross Abstract: Despite extensive alignment efforts, Large Vision-Language Models (LVLMs) remain vulnerable to jailbreak attacks. To mitigate these risks, existing detection methods are essential, yet they face two major challenges: gener…
MMJailBench: prompt framing is top jailbreak risk across 16 AI models MMJailBench, a new factorized benchmark, evaluated 16 multimodal LLMs and found prompt framing, not visual tricks, drives most jailbreak vulnerabilities. https://www. notatechguy.com/mmjailbench-pr ompt-framing…