Researchers have developed a novel attack that can reconstruct text generated by locally hosted Large Language Models (LLMs) by observing CPU cache activity during the detokenization process. This method, termed "Detokenization Leaks," bypasses previous attack limitations by targeting the detokenizer, a common component in LLM inference pipelines. By using Flush+Reload and Prime+Probe techniques to monitor cache behavior, the attack can recover semantically accurate outputs from various LLM deployments, including agentic systems, across different hardware and software configurations. AI
IMPACT Highlights a new class of side-channel attacks against local LLMs, potentially impacting the security of sensitive data processed by these models.
RANK_REASON Academic paper detailing a new security vulnerability in LLM detokenization. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- CatalyzeX
- CPU cache
- DagsHub
- Detokenization Leaks
- Flush+Reload
- Hugging Face
- LLM
- OpenClaw
- Prime+Probe
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →