PulseAugur
EN
LIVE 11:39:24

Kimi K3 lags US models in cyber exploits, possible distillation cited

Moonshot AI's Kimi K3 model demonstrated significantly weaker performance on offensive cyber tasks compared to leading U.S. models, scoring 32% on ExploitBench versus 76%. The model's safeguards also failed to prevent exploit development or simulated attacks. This disparity between Kimi K3's general benchmark scores and its cyber capabilities aligns with claims that Moonshot AI may have used distillation techniques, potentially involving Anthropic's models. AI

IMPACT Highlights potential vulnerabilities in AI models used for cyber security and suggests that distillation techniques may impact specialized capabilities.

RANK_REASON Research report on AI model performance in cyber security tasks. [lever_c_demoted from research: ic=1 ai=1.0]

Read on The Decoder →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Kimi K3 lags US models in cyber exploits, possible distillation cited

COVERAGE [1]

  1. The Decoder TIER_1 English(EN) · Matthias Bastian ·

    Kimi K3 trails frontier US models by a wide margin on cyber exploits, and distillation may explain why

    <p><img alt="" class="attachment-full size-full wp-post-image" height="1152" src="https://the-decoder.com/wp-content/uploads/2026/07/aisi_logo_pattern.png" style="height: auto; margin-bottom: 10px;" width="2048" /></p> <p> The British AI Security Institute and the U.S. Center for…