Researchers have discovered a method to extract hidden Chain-of-Thought (CoT) reasoning from frontier AI models by exploiting a vulnerability in their APIs. This vulnerability allows for the retrieval of internal reasoning processes when a "deep_think" tool is provided instead of direct thinking capabilities. The extracted reasoning token count precisely matches the billed API thinking tokens for most queried prompts, indicating a potential issue with how AI companies are accounting for internal model computations. AI
IMPACT This finding could pressure AI labs to improve API security and transparency regarding internal model computations.
RANK_REASON Researchers detail a method to extract hidden reasoning from AI models. [lever_c_demoted from research: ic=1 ai=1.0]
Read on HN — anthropic stories →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →