Researchers have developed a new type of backdoor attack targeting large language models (LLMs) that operates on the answer side, rather than the input side. This novel approach involves a benign initial prompt that causes the LLM to generate a specific word, which then acts as a trigger when embedded in the dialogue history. When a subsequent harmful query is made, the model recognizes its own generated trigger and bypasses safety protocols, even though the user's input remains clean. This answer-side backdoor achieved near-perfect attack success rates across multiple LLMs with a low poisoning rate, evading current input-centric defenses and highlighting a significant vulnerability in LLM safety alignment. AI
IMPACT Highlights a critical blind spot in current LLM safety alignment, potentially requiring new defense strategies.
RANK_REASON Academic paper detailing a novel attack vector against LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
- answer-side backdoor
- arXiv
- Attack Success Rates
- dialogue history
- harmful query
- input-centric defenses
- large-language models
- LLM defenses
- Multi-turn dialogue response generation via mutual information maximization
- poisoning rate
- Representation-level analysis
- User input interfaces
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →