A new paper reveals a vulnerability in proprietary LLM APIs from Anthropic, OpenAI, and Google, allowing encrypted reasoning traces to be extracted and replayed. Researchers demonstrated that these encrypted "thought blocks" are interchangeable across models and sessions within the same provider's ecosystem. By injecting an encrypted trace into a weaker model, attackers can force it to reveal the stronger model's hidden reasoning, bypassing anti-distillation measures. This technique can also expose sensitive data like API keys and personal information that were inadvertently included in these hidden traces, and potentially enable invisible prompt injections. AI
IMPACT Exposes sensitive data and intellectual property, potentially impacting enterprise adoption and security practices for LLM APIs.
RANK_REASON Paper detailing a vulnerability in LLM API reasoning trace encryption.
Read on Hugging Face Daily Papers →
- Alexander V Panfilov
- Anthropic
- OpenAI
- Claude Haiku 4.5
- GPT-5.5
- GPT 5.6 Luna
- proprietary LLM APIs
- Simon Willison
- arXiv
- Claude (Opus 4.8)
- Hacker News
- Matthew D. Green
AI-generated summary · Google Gemini · from 8 sources. How we write summaries →