ExploitBench
PulseAugur coverage of ExploitBench — every cluster mentioning ExploitBench across labs, papers, and developer communities, ranked by signal.
- 2026-06-18 research_milestone Carnegie Mellon University researchers developed ExploitBench to measure AI model exploit capabilities. source
3 day(s) with sentiment data
-
Zhipu AI's GLM-5.3 shows mixed results in cybersecurity benchmarks
Zhipu AI has released its GLM-5.3 model, which has generated headlines for its purported superior performance in cybersecurity benchmarks. While the model did achieve a slightly higher score than Anthropic's Mythos 5 an…
-
Zhipu AI's GLM-5.3 excels in coding and cybersecurity benchmarks · 8 sources tracked
Zhipu AI has launched GLM-5.3, an open-weights coding model that reportedly shows significant improvements over its predecessor, GLM-5.2, primarily through scaled post-training. The model demonstrates notable gains in c…
-
UK/US assess Kimi K3 cyber capabilities, finding it lags frontier models
A preliminary assessment by the UK Artificial Intelligence Safety Institute (UK AISI) and the U.S. Center for AI Standards and Innovation (CAISI) has evaluated the cybersecurity capabilities of Moonshot AI's Kimi K3 mod…
-
Kimi K3 lags US models in cyber exploit tests, distillation suspected · 2 sources tracked
Moonshot AI's Kimi K3 model demonstrated significantly weaker performance on cyber exploit tasks compared to leading US models, scoring 32% on ExploitBench versus 76%. The model's safeguards also proved insufficient in …
-
China's Kimi K3 lags US rivals in cybersecurity capabilities, study finds
A joint UK-US government study found that China's Kimi K3 large language model significantly underperforms leading US models in cybersecurity capabilities. The research, conducted using the ExploitBench benchmark, showe…
-
Apple bolsters security after supplier breach and AI hacking fears
Apple is enhancing its cybersecurity measures in response to a significant supply chain attack on its supplier Tata, which resulted in leaked client files and information about an unreleased iPhone model. The company is…
-
AI models tested for exploit capabilities; Anthropic's Mythos shows advanced execution
Researchers at Carnegie Mellon University have developed ExploitBench, a new framework to measure how effectively AI models can exploit security vulnerabilities. While most public frontier models cause crashes, they gen…
-
Anthropic's powerful Claude Mythos AI breached via contractor access
Anthropic's highly capable cybersecurity AI model, Claude Mythos, was reportedly accessed by unauthorized users shortly after its limited preview began. The breach occurred through a combination of insider knowledge fro…