Anthropic's Frontier Red Team has identified that advanced AI models are now capable of developing full control flow hijacks in binary exploitation tasks. GLM-5.3 achieved this in 4% of trials, while Claude Mythos Preview did so in 6%. This represents a significant leap, as earlier models like Claude Opus 4.6 and GLM-5.2 were unable to perform such exploits. AI
IMPACT Emerging AI models demonstrate advanced cyber exploit capabilities, raising concerns for cybersecurity defenses.
RANK_REASON The cluster reports on research findings from an AI red team regarding model capabilities in cybersecurity exploits.
- Anthropic
- Anthropic Frontier Red Team
- Claude Mythos Preview
- Claude Opus 4.6
- Claude Opus 5.5
- GLM-5.2
- GLM-5.3
- GPT-6 Luna
- GPT-6 Sol
- OpenAI DevDay 2026
- Simon Willison
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →