Anthropic has detailed sophisticated distillation attacks originating from Chinese AI companies, including Alibaba Group, Moonshot AI, and DeepSeek. These attacks aim to extract proprietary reasoning capabilities from Anthropic's Claude models to train smaller, competing models. The company observed nearly 200 million exchanges linked to these attacks, with one campaign attributed to Alibaba being the largest wholesale distillation effort ever recorded. Moonshot AI's campaign, in particular, appeared to route requests from the Chinese military, targeting Claude's Opus model. AI
IMPACT Highlights the escalating security challenges in frontier model development and the potential for intellectual property theft.
RANK_REASON Report detailing AI model security vulnerabilities and adversarial attacks.
Read on Mastodon — mastodon.social →
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →