Researchers have identified a new class of vulnerabilities in large language models (LLMs) called Constrained Decoding Attack (CDA). This attack targets the control plane of LLMs, exploiting the grammar-guided decoding process to inject malicious content. Unlike traditional data-plane jailbreaks, CDA operates on the decoding mechanism itself, making it difficult for internal safety alignments to prevent. A specific instantiation, DictAttack, achieved a 94.3-99.5% attack success rate on prominent models like GPT-5 and Gemini 2.5 Pro, even bypassing state-of-the-art jailbreak defenses. AI
IMPACT Highlights a critical control-plane vulnerability in LLMs, potentially requiring new defense mechanisms beyond current safety alignments.
RANK_REASON Academic paper detailing a new class of LLM vulnerabilities. [lever_c_demoted from research: ic=1 ai=1.0]
- Constrained Decoding Attack (CDA)
- DeepSeek-R1
- DictAttack
- EnumAttack
- Gemini 2.5 Pro
- GPT-5
- GPT-OSS 120B
- Shuoming Zhang
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →