English(EN)SALLIE: Generation-Free Hidden-State Detection of Jailbreaks and Prompt Injections Across Text and Vision
新研究通过先进的检测和防御策略应对LLM越狱 · 跟踪7个来源
作者PulseAugur 编辑部·[8 个来源]·
研究人员正在开发先进的方法来检测和防止针对大型语言和视觉语言模型的越狱攻击。SALLIE等新技术通过分析模型内部状态提供免生成、跨模态检测,而NeuroBreak则专注于理解神经元级别的安全机制。其他方法包括使用心理学原理自动生成多轮攻击数据集,以及利用已知攻击示例的检索增强防御(RAD)等自适应防御。新颖的策略还探讨了诱饵图像如何放大现有防御能力,以及迭代上下文优化如何增强语义偏移越狱。
AI
arXiv:2608.08542v1 Announce Type: new Abstract: Model merging has become the default way to give an aligned language model new skills without retraining: a practitioner folds task vectors from math, code, or domain specialists into a safety-aligned base using task arithmetic, TIE…
arXiv cs.AI
TIER_1Deutsch(DE)·Chuhan Zhang, Ye Zhang, Bowen Shi, Yuyou Gan, Tianyu Du, Shouling Ji, Dazhan Deng, Yingcai Wu·
arXiv:2509.03985v2 Announce Type: replace-cross Abstract: In deployment and application, large language models (LLMs) typically undergo safety alignment to prevent illegal and unethical outputs. However, the continuous advancement of jailbreak attack techniques, designed to bypas…
arXiv:2511.19517v3 Announce Type: replace-cross Abstract: Multi-turn conversational attacks, which leverage psychological principles like Foot-in-the-Door (FITD), where a small initial request paves the way for a more significant one, to bypass safety alignments, pose a persisten…
arXiv cs.AI
TIER_1English(EN)·Guy Azov, Ofer Rivlin, Guy Shtar·
arXiv:2604.06247v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) and Vision-Language Models (VLMs) are vulnerable to jailbreaks and prompt injections delivered through text or images. Existing defenses often narrow threat coverage or add inference cost throu…
arXiv:2608.01043v2 Announce Type: replace-cross Abstract: We report a counter-intuitive interaction between image inputs and existing black-box defenses on Vision--Language Models (VLMs): pairing an encoded jailbreak prompt with an unrelated decoy image can sharply lower attack s…
arXiv:2508.16406v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) remain vulnerable to jailbreak attacks, which attempt to elicit harmful responses from LLMs. The evolving nature and diversity of these attacks pose many challenges for defense systems, includi…
Foundation models have achieved remarkable success across diverse tasks, but they remain vulnerable. To investigate such vulnerabilities, semantic-shift jailbreaks have recently emerged as a promising attack paradigm. They bypass explicit safety mechanisms by replacing harmful te…
arXiv:2608.10933v1 Announce Type: new Abstract: Text-to-Video (T2V) generative models are vulnerable to jailbreak attacks in real-world deployment, leading them to produce harmful or inappropriate content. Existing defense approaches mainly rely on input filtering or reconstructi…