PulseAugur
中
实时 18:43:08
English(EN) Fully Unleashing the Multimodal Attacker: Meta-Adaptive Jailbreaking of Vision-Language Models

新研究探讨了大型语言模型和视觉-语言模型的先进越狱技术及检测方法

研究人员正在开发先进的方法来测试大型语言模型和视觉-语言模型在面对越狱尝试时的安全性和鲁棒性。SEAV等新框架专注于验证模型响应的正确性和程序准确性,而不仅仅是语义上的合理性。其他研究引入了元自适应攻击,该攻击会优化攻击者本身以利用多模态模型的漏洞,在GPT-4o和Gemini 3 Pro Preview等领先模型上取得了高成功率。此外,还在探索通过分析模型激活来检测未见过越狱攻击的方法,旨在提高安全评估的泛化能力和效率。 AI

影响 这些越狱和检测方面的进展凸显了持续的安全挑战,并推动了AI模型更强大防御措施的开发。

排序理由 该集群包含多篇研究论文,详细介绍了用于越狱和检测大型语言模型及视觉-语言模型漏洞的新颖方法。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 10 个来源。 我们如何撰写摘要 →

新研究探讨了大型语言模型和视觉-语言模型的先进越狱技术及检测方法

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群包含多篇研究论文,详细介绍了用于越狱和检测大型语言模型及视觉-语言模型漏洞的新颖方法。
Source corroboration
10 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
safety, paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
49 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+2 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [10]

  1. arXiv cs.AI TIER_1 English(EN) · Saikat Mondal, Mamta, Deeksha Varshney, Oana Cocarascu, Asif Ekbal ·

    IndicSafeEval:大型语言模型在多语言说服性越狱攻击下的安全鲁棒性

    arXiv:2609.03781v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used in multilingual settings, yet their safety is still evaluated primarily in English. This limits our understanding of how alignment failures manifest in low-resource and culturally…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    IndicSafeEval:多语言说服性越狱攻击下大语言模型的安全鲁棒性

    Large language models (LLMs) are increasingly used in multilingual settings, yet their safety is still evaluated primarily in English. This limits our understanding of how alignment failures manifest in low-resource and culturally diverse languages. We introduce IndicSafeEval, a …

  3. arXiv cs.AI TIER_1 English(EN) · Qilong Wu, Sahil Wadhwa, Pranab Mohanty, Giri Iyengar, Varun Chandrasekaran ·

    面向大型语言模型的有效性感知越狱评估

    arXiv:2609.00498v1 Announce Type: new Abstract: Jailbreak robustness has become central to large language model (LLM) safety evaluation, yet prevailing methodologies rely primarily on refusal behavior, semantic resemblance, and intent-matching heuristics that emphasize linguistic…

  4. Hugging Face Daily Papers TIER_1 English(EN) ·

    全面释放多模态攻击者:视觉语言模型的元自适应越狱

    The safety of large vision-language models is increasingly stress-tested by multimodal jailbreaks, yet existing attacks remain largely static at the meta level: template-based attacks freeze the image--text layout, while iterative attacks adapt only the image--text content with f…

  5. Hugging Face Daily Papers TIER_1 English(EN) ·

    面向多模态大语言模型的、可迁移越狱攻击的文本锚定语义扰动

    Multimodal Large Language Models (MLLMs) have achieved remarkable progress in vision-language interaction, yet their safety alignment remains vulnerable to jailbreak attacks. A key challenge is that safety behavior learned in the textual space does not reliably transfer to fused …

  6. Hugging Face Daily Papers TIER_1 English(EN) ·

    TempJail:针对大型视觉语言模型的字幕调度时间性越狱攻击

    Large vision-language models (LVLMs) have achieved remarkable progress in video understanding and reasoning. Despite extensive studies on text- and image-based jailbreaks, video jailbreaks against LVLMs remain largely unexplored. Existing video jailbreak methods mainly manipulate…

  7. arXiv cs.CV TIER_1 English(EN) · Benlei Cui, Shen Pang, Yuke Wang, Xuemei Dong, Yuwen Zhai, Jingqun Tang, Haiyang Yu, Hui Xue, Longtao Huang, Haiwen Hong ·

    全面释放多模态攻击者:视觉语言模型的元自适应越狱

    arXiv:2608.27531v1 Announce Type: cross Abstract: The safety of large vision-language models is increasingly stress-tested by multimodal jailbreaks, yet existing attacks remain largely static at the meta level: template-based attacks freeze the image--text layout, while iterative…

  8. arXiv cs.CV TIER_1 English(EN) · Qi Lu, Zehui Guo, David Yuanda Gan, Zijing Li, Hengda Zhang, Weijun Xu, Qiankun Zhang ·

    TempJail:针对图像到视频生成模型的时序越狱攻击

    arXiv:2608.26971v1 Announce Type: new Abstract: In recent years, image-to-video (I2V) generation models have made remarkable progress in subject consistency and temporal coherence, enabling high quality video synthesis. However, these advances also introduce new safety risks. Exi…

  9. arXiv cs.CV TIER_1 English(EN) · Shuang Liang, Zhihao Xu, Jiaqi Weng, Jialing Tao, Hui Xue, Xiting Wang ·

    学习检测大型视觉语言模型中未见的越狱攻击

    arXiv:2508.09201v5 Announce Type: replace-cross Abstract: Despite extensive alignment efforts, Large Vision-Language Models (LVLMs) remain vulnerable to jailbreak attacks. To mitigate these risks, existing detection methods are essential, yet they face two major challenges: gener…

  10. Mastodon — mastodon.social TIER_1 English(EN) · notatechguy ·

    MMJailBench:提示框架是 16 个 AI 模型中的首要越狱风险 MMJailBench,一个新颖的因子化基准测试,评估了 16 个多模态大语言模型,并发现提示框架

    MMJailBench: prompt framing is top jailbreak risk across 16 AI models MMJailBench, a new factorized benchmark, evaluated 16 multimodal LLMs and found prompt framing, not visual tricks, drives most jailbreak vulnerabilities. https://www. notatechguy.com/mmjailbench-pr ompt-framing…