PulseAugur
实时 08:25:16
English(EN) SALLIE: Generation-Free Hidden-State Detection of Jailbreaks and Prompt Injections Across Text and Vision

新研究通过先进的检测和防御策略应对LLM越狱 · 跟踪7个来源

研究人员正在开发先进的方法来检测和防止针对大型语言和视觉语言模型的越狱攻击。SALLIE等新技术通过分析模型内部状态提供免生成、跨模态检测,而NeuroBreak则专注于理解神经元级别的安全机制。其他方法包括使用心理学原理自动生成多轮攻击数据集,以及利用已知攻击示例的检索增强防御(RAD)等自适应防御。新颖的策略还探讨了诱饵图像如何放大现有防御能力,以及迭代上下文优化如何增强语义偏移越狱。 AI

影响 越狱检测和防御的进步对于提高LLM和VLM在实际应用中的安全性和可靠性至关重要。

排序理由 多篇研究论文详细介绍了检测和防御LLM和VLM越狱攻击的新方法。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 8 个来源。 我们如何撰写摘要 →

新研究通过先进的检测和防御策略应对LLM越狱 · 跟踪7个来源

报道来源 [8]

  1. arXiv cs.LG TIER_1 English(EN) · Yu Ma, Hongli Shi, Jing Li, Xinran Xu, Weiwei Hou ·

    技能与安全交汇:技能融合大模型的自适应越狱鲁棒性基准测试与表征

    arXiv:2608.08542v1 Announce Type: new Abstract: Model merging has become the default way to give an aligned language model new skills without retraining: a practitioner folds task vectors from math, code, or domain specialists into a safety-aligned base using task arithmetic, TIE…

  2. arXiv cs.AI TIER_1 Deutsch(DE) · Chuhan Zhang, Ye Zhang, Bowen Shi, Yuyou Gan, Tianyu Du, Shouling Ji, Dazhan Deng, Yingcai Wu ·

    NeuroBreak:揭示大型语言模型内部越狱机制

    arXiv:2509.03985v2 Announce Type: replace-cross Abstract: In deployment and application, large language models (LLMs) typically undergo safety alignment to prevent illegal and unethical outputs. However, the continuous advancement of jailbreak attack techniques, designed to bypas…

  3. arXiv cs.AI TIER_1 English(EN) · Adarsh Kumarappan, Ananya Mujoo ·

    自动化欺骗:可扩展的多轮 LLM 越狱

    arXiv:2511.19517v3 Announce Type: replace-cross Abstract: Multi-turn conversational attacks, which leverage psychological principles like Foot-in-the-Door (FITD), where a small initial request paves the way for a more significant one, to bypass safety alignments, pose a persisten…

  4. arXiv cs.AI TIER_1 English(EN) · Guy Azov, Ofer Rivlin, Guy Shtar ·

    SALLIE:跨越文本和视觉的无生成隐藏状态越狱和提示注入检测

    arXiv:2604.06247v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) and Vision-Language Models (VLMs) are vulnerable to jailbreaks and prompt injections delivered through text or images. Existing defenses often narrow threat coverage or add inference cost throu…

  5. arXiv cs.AI TIER_1 English(EN) · Haoyu Zhang, Xiangchen Guan, Shibo Zheng, Mohammad Zandsalimy, Shanu Sushmita ·

    诱饵图像增强了针对编码越狱的字幕介导防御

    arXiv:2608.01043v2 Announce Type: replace-cross Abstract: We report a counter-intuitive interaction between image inputs and existing black-box defenses on Vision--Language Models (VLMs): pairing an encoded jailbreak prompt with an unrelated decoy image can sharply lower attack s…

  6. arXiv cs.CL TIER_1 English(EN) · Guangyu Yang, Jinghong Chen, Jingbiao Mei, Weizhe Lin, Bill Byrne ·

    检索增强防御:大型语言模型自适应可控的越狱防护

    arXiv:2508.16406v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) remain vulnerable to jailbreak attacks, which attempt to elicit harmful responses from LLMs. The evolving nature and diversity of these attacks pose many challenges for defense systems, includi…

  7. Hugging Face Daily Papers TIER_1 English(EN) ·

    ICO:通过迭代上下文优化增强语义偏移越狱

    Foundation models have achieved remarkable success across diverse tasks, but they remain vulnerable. To investigate such vulnerabilities, semantic-shift jailbreaks have recently emerged as a promising attack paradigm. They bypass explicit safety mechanisms by replacing harmful te…

  8. arXiv cs.CV TIER_1 English(EN) · Siyuan Liang, Yupeng Qiu, Junfeng Fang, Rong-Cheng Tu, Jiaxing Huang, Dacheng Tao ·

    SafeCA:文本到视频越狱防御的安全交叉注意力定位与调控

    arXiv:2608.10933v1 Announce Type: new Abstract: Text-to-Video (T2V) generative models are vulnerable to jailbreak attacks in real-world deployment, leading them to produce harmful or inappropriate content. Existing defense approaches mainly rely on input filtering or reconstructi…