PulseAugur
实时 06:48:40
English(EN) Jailbreaking Text-to-Image Models Through Cracks: Navigating Heterogeneous Safety Filters via Multi-Agent Debate

新的 CRACK 框架绕过了文本到图像模型的安全过滤器

研究人员开发了一个名为 CRACK 的新颖多智能体辩论框架,用于绕过文本到图像模型的安全过滤器。该框架使用攻击智能体、防御智能体和裁判智能体来迭代地优化提示,识别不同安全层之间的冲突,并保持原始的有害意图。CRACK 在生成不适合工作场所的内容方面表现出很高的成功率,即使在复杂的安全配置下也是如此,同时所需的查询次数比现有方法少,并保持了语义保真度。 AI

影响 这项研究突显了当前文本到图像安全机制的漏洞,可能会推动更强大的防御措施的开发。

排序理由 该集群包含一篇详细介绍越狱 AI 模型新方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的 CRACK 框架绕过了文本到图像模型的安全过滤器

本文如何被排名

Signal score
27 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群包含一篇详细介绍越狱 AI 模型新方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Kaiyan Wen, Shijie Zhang, Lu Yu, Guangdong Bai ·

    通过裂缝越狱文本到图像模型:通过多智能体辩论导航异构安全过滤器

    arXiv:2609.01168v1 Announce Type: new Abstract: Text-to-image (T2I) models remain vulnerable to jailbreak attacks that elicit Not-Safe-For-Work (NSFW) content, despite increasingly being guarded by heterogeneous, multi-layer safety stacks combining text filters, image classifiers…