PulseAugur
EN
LIVE 04:50:00

Researchers exploit API vulnerability to extract hidden AI model reasoning

Researchers have discovered a method to extract hidden Chain-of-Thought (CoT) reasoning from frontier AI models by exploiting a vulnerability in their APIs. This vulnerability allows for the retrieval of internal reasoning processes when a "deep_think" tool is provided instead of direct thinking capabilities. The extracted reasoning token count precisely matches the billed API thinking tokens for most queried prompts, indicating a potential issue with how AI companies are accounting for internal model computations. AI

IMPACT This finding could pressure AI labs to improve API security and transparency regarding internal model computations.

RANK_REASON Researchers detail a method to extract hidden reasoning from AI models. [lever_c_demoted from research: ic=1 ai=1.0]

Read on HN — anthropic stories →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Researchers exploit API vulnerability to extract hidden AI model reasoning

COVERAGE [1]

  1. HN — anthropic stories TIER_1 English(EN) · himata4113 ·

    OpenAI and Anthropic hidden CoT leaks when given deep_think tool.