Researchers have developed a novel zero-query jailbreak framework designed to bypass safety filters in text-to-image systems. This method exploits the Filter-Generator Discrepancy (FGD), where prompt perturbations are processed differently by the safety filter and the image generator. By identifying observable discrepancies at tokenization and semantic levels, the framework screens for high-potential candidates before employing an ensemble evolutionary search that requires no direct access to the target system. Experiments demonstrated a significant increase in attack success rates, outperforming existing baselines on multiple pipelines and a commercial service. AI
IMPACT This research highlights potential vulnerabilities in AI safety filters, suggesting a need for more robust defenses against adversarial attacks in generative AI systems.
RANK_REASON Academic paper detailing a new method for jailbreaking AI systems. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Canada
- Filter-Generator Discrepancy
- Hugging Face
- Master of Health Science
- Mind the Gap: Zero-Query Jailbreaks via Filter-Generator Discrepancy in Text-to-Image Systems
- Text-to-Image Systems
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →