Researchers have discovered a vulnerability in the API ecosystems of Anthropic, OpenAI, and Google that allows for the extraction of supposedly hidden reasoning traces. By replaying encrypted reasoning blocks into weaker compatible models, these models act as decryption oracles, revealing the opaque traces as readable plaintext. A scan of public agent trajectories uncovered personal information and credentials within these decoded blocks, highlighting a potential security risk for sensitive data embedded in AI reasoning outputs. AI
IMPACT Highlights a potential security flaw in how AI reasoning data is handled, urging caution for sensitive information.
RANK_REASON Academic paper detailing a novel attack vector on AI reasoning outputs. [lever_c_demoted from research: ic=1 ai=1.0]
Read on dev.to — Anthropic tag →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →