Researchers have developed a novel backdoor attack called Paraesthesia that targets large language models (LLMs) by leveraging emotional style in inputs. Unlike previous attacks that rely on fixed triggers, Paraesthesia encodes its malicious condition within the emotional tone of text, achieving over 98.25% attack success rate across various tasks and LLMs. This method demonstrates that emotional style can serve as a trigger surface for backdoors, distinct from traditional lexical or syntactic patterns, and proves resilient to several defense mechanisms. AI
IMPACT Identifies a new attack vector for LLMs, potentially impacting model security and the effectiveness of current defense strategies.
RANK_REASON Academic paper detailing a new method for attacking LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →