A joint UK-US government study found that China's Kimi K3 large language model significantly underperforms leading US models in cybersecurity capabilities. The research, conducted using the ExploitBench benchmark, showed Kimi K3 scoring 32.2% overall, while top US models averaged 76.2%. Kimi K3 was unable to achieve arbitrary code execution, a critical exploit capability that top US models demonstrated on multiple tasks. AI
IMPACT Highlights potential cybersecurity risks associated with Chinese AI models compared to US counterparts.
RANK_REASON Research report on AI model capabilities using a benchmark. [lever_c_demoted from research: ic=1 ai=1.0]
- ExploitBench
- GLM-5.2
- Kimi k3
- Moonshot AI
- National Institute of Standards and Technology
- UK Artificial Intelligence Security Institute
- UK Department for Science, Innovation and Technology
- United States Department of Commerce
- US Centre for AI Standards and Innovation
- Zhipu AI
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →