Researchers have discovered a vulnerability in how major LLM providers like Anthropic, OpenAI, and Google handle AI reasoning traces. These providers encrypt reasoning steps and send them back to the client, but the encrypted blocks are interchangeable across sessions and models. Attackers can exploit this to decrypt proprietary reasoning, extract private data including PII and credentials, reveal hazardous information, and perform invisible prompt injections. The researchers propose cryptographic and system-level mitigations. AI
IMPACT Reveals a new attack vector for extracting proprietary LLM logic and sensitive data, potentially impacting model security and user privacy.
RANK_REASON Research paper detailing a vulnerability in LLM reasoning trace handling.
Read on Mastodon — mastodon.social →
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →