PulseAugur
实时 08:29:44
English(EN) Robustness of Vision Language Models Against Split-Image Harmful Input Attacks

新发现的 VLM 在分片图像攻击中存在漏洞

研究人员发现了一种新的 Vision-Language Models (VLMs) 漏洞,即安全对齐功能无法有效抵御被分割到多个图像片段中的有害内容。虽然 VLMs 在预训练过程中能够很好地泛化到分片图像,但它们的安全机制(通常在整体图像上训练)却无法检测到组合起来的有害语义。该研究引入了新颖的分片图像视觉越狱攻击 (SIVA),这些攻击从简单的分割逐渐演变为自适应的白盒和黑盒迁移攻击,并利用对抗性知识蒸馏 (Adv-KD) 来增强跨模型的可迁移性。评估结果表明,与现有方法相比,这些攻击在最先进的 VLMs 上取得了显著更高的成功率。 AI

影响 凸显了 VLM 安全对齐方面的一个关键差距,可能需要新的训练方法来应对分片图像攻击。

排序理由 详细介绍 VLM 新漏洞和攻击方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新发现的 VLM 在分片图像攻击中存在漏洞

本文如何被排名

Signal score
17 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
详细介绍 VLM 新漏洞和攻击方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Md Rafi Ur Rashid, MD Sadik Hossain Shanto, Vishnu Asutosh Dasu, Shagufta Mehnaz ·

    Vision Language Models Against Split-Image Harmful Input Attacks 的鲁棒性

    arXiv:2602.08136v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) are now a core part of modern AI. Recent work proposed several visual jailbreak attacks using single/ holistic images. However, contemporary VLMs demonstrate strong robustness against such att…