Researchers have discovered a vulnerability in how major LLM providers like Anthropic, OpenAI, and Google handle proprietary model reasoning traces. By exploiting the interchangeability of encrypted reasoning blocks across sessions and models, attackers can force weaker models to reveal plaintext traces from more capable ones. This method allows for the extraction of proprietary reasoning, the retrieval of sensitive data like PII and credentials from public logs, the exposure of hazardous information, and the execution of invisible prompt injections. Proposed mitigations focus on cryptographic and system-level security enhancements. AI
IMPACT Exposes a critical security flaw in LLM APIs, potentially impacting data privacy, intellectual property, and the security of agentic systems.
RANK_REASON Academic paper detailing a novel vulnerability and attack vectors in LLM API security. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →