Researchers have developed a Persistent Memory Poisoning Attack (PMPA) targeting LLM-based agents that utilize memory, tool use, and runtime control. This attack embeds malicious instructions into benign external sources, tricking agents into storing them in persistent memory. Once stored, these poisoned instructions can be retrieved in subsequent sessions, leading to unintended malicious actions and privacy breaches. Evaluations on OpenClaw and Claude Code demonstrated significant success rates for injecting and executing these malicious instructions across various configurations, with defenses showing limited effectiveness against already poisoned memory. AI
IMPACT Highlights critical security vulnerabilities in LLM agent architectures, necessitating robust defense mechanisms.
RANK_REASON Academic paper detailing a new attack vector on LLM agents. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →