PulseAugur
EN
LIVE 20:38:37
ENTITY ExploitBench

ExploitBench

PulseAugur coverage of ExploitBench — every cluster mentioning ExploitBench across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
21
21 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
2
2 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
TIMELINE
  1. 2026-06-18 research_milestone Carnegie Mellon University researchers developed ExploitBench to measure AI model exploit capabilities. source
SENTIMENT · 30D

6 day(s) with sentiment data

RECENT · PAGE 1/2 · 21 TOTAL
  1. TOOL · CL_276491 ·

    GLM 5.3 nears Anthropic's Mythos Preview on ExploitBench, with safety removal cost estimated at $1,200

    An open-weight model named GLM 5.3 has demonstrated performance approaching that of Anthropic's restricted Mythos Preview on the ExploitBench benchmark. While GLM 5.3's capabilities are nearing frontier models, estimate…

  2. RESEARCH · CL_269032 ·

    Anthropic flags Zhipu AI's GLM-5.3 for advanced cyber exploit capabilities

    Anthropic has released findings on GLM-5.3, a new AI model from Zhipu AI, highlighting its advanced capabilities in autonomously building cyber exploits. Unlike Anthropic's own Claude Mythos Preview, which was released …

  3. COMMENTARY · CL_248990 ·

    AGI debate heats up with new benchmarks and OpenAI's Astra model

    A recent video and accompanying research explore the evolving definition and potential achievement of Artificial General Intelligence (AGI). The discussion contrasts economic definitions of AGI with frameworks focusing …

  4. COMMENTARY · CL_241843 ·

    OpenAI's GPT-6 Astra benchmark scores questioned; alternative testing proposed

    A recent analysis of OpenAI's new GPT-6 Astra model highlights potential issues with its benchmark scores. While OpenAI reported near-perfect results on several tests, including ExploitBench, the author points out that …

  5. RESEARCH · CL_240147 ·

    GPT 6 achieves perfect score on ExploitBench, raising concerns for Hugging Face

    GPT 6 has achieved a perfect 100% score on the ExploitBench benchmark, a feat that raises concerns for Hugging Face. The report suggests that Hugging Face may not have disclosed a breach related to this achievement. Thi…

  6. SIGNIFICANT · CL_235828 ·

    OpenAI's GPT-6 Astra classified 'Critical' for autonomous cyber exploit discovery

    OpenAI has released GPT-6 Astra, its first model classified as "Critical" in cybersecurity preparedness due to its ability to autonomously discover and exploit zero-day vulnerabilities in hardened systems. Astra achieve…

  7. COMMENTARY · CL_235174 ·

    Open-weight AI models challenge frontier models, closing performance gap

    Open-weight AI models are rapidly closing the performance gap with closed frontier models, with some Chinese models now rivaling top US offerings in benchmarks. While Kimi K3 from Moonshot AI leads open-weight models, i…

  8. FRONTIER RELEASE · CL_231906 ·

    GPT-6 Astra benchmarks show leap over Fable 5.1, users anticipate OpenAI advancement

    Users are discussing the upcoming GPT-6 Astra model, with early benchmarks suggesting it outperforms Fable 5.1 and GPT Sol. GPT-6 Astra is noted for its high scores on various benchmarks, including ARC-AGI-3 and SRE-Ben…

  9. FRONTIER RELEASE · CL_230959 ·

    OpenAI launches GPT-6 Astra, declares 'AGI era' amid performance debates · 10 sources tracked

    OpenAI has officially launched GPT-6 Astra and GPT-6 Astra Pro, its most advanced models to date, signaling a potential entry into the era of Artificial General Intelligence (AGI). These models demonstrate significant a…

  10. SIGNIFICANT · CL_230838 ·

    OpenAI's Astra model nears release with advanced cybersecurity capabilities

    OpenAI is preparing to release its new Astra model, which it claims is the first large language model to meet its stringent cybersecurity threshold. Astra has demonstrated a remarkable ability to identify and exploit un…

  11. FRONTIER RELEASE · CL_230849 ·

    OpenAI releases GPT-6 Astra for business with advanced reasoning

    OpenAI has announced GPT-6 Astra, its latest model designed for business applications. This new model boasts enhanced reasoning capabilities, improved computer use, and better judgment in writing and design. GPT-6 Astra…

  12. FRONTIER RELEASE · CL_229672 ·

    OpenAI releases GPT-6 Astra, touting advanced capabilities and safety

    OpenAI has released its latest model, GPT-6 Astra, which is now available to users across various tiers including Pro, Enterprise, and Business Premium, as well as through its API and on Amazon Bedrock. This model is de…

  13. RESEARCH · CL_207346 ·

    AI models find bugs and bypass safety filters, fueling security incidents

    A new frontier coding model, GLM 5.3, discovered a live vulnerability in the AI-powered code editor Cursor within a day of its release. This highlights a growing concern where the same AI models capable of identifying s…

  14. TOOL · CL_206946 ·

    Zhipu AI's GLM-5.3 shows mixed results in cybersecurity benchmarks

    Zhipu AI has released its GLM-5.3 model, which has generated headlines for its purported superior performance in cybersecurity benchmarks. While the model did achieve a slightly higher score than Anthropic's Mythos 5 an…

  15. FRONTIER RELEASE · CL_200360 ·

    Z.ai's GLM-5.3 achieves SOTA in coding and cybersecurity via post-training

    Z.ai has released GLM-5.3, an updated model that achieves significant performance gains through scaled post-training rather than changes to its base model. The model shows marked improvements in complex coding tasks, pa…

  16. RESEARCH · CL_161580 ·

    UK/US assess Kimi K3 cyber capabilities, finding it lags frontier models

    A preliminary assessment by the UK Artificial Intelligence Safety Institute (UK AISI) and the U.S. Center for AI Standards and Innovation (CAISI) has evaluated the cybersecurity capabilities of Moonshot AI's Kimi K3 mod…

  17. RESEARCH · CL_161358 ·

    Kimi K3 lags US models in cyber exploit tests, distillation suspected · 2 sources tracked

    Moonshot AI's Kimi K3 model demonstrated significantly weaker performance on cyber exploit tasks compared to leading US models, scoring 32% on ExploitBench versus 76%. The model's safeguards also proved insufficient in …

  18. TOOL · CL_161121 ·

    China's Kimi K3 lags US rivals in cybersecurity capabilities, study finds

    A joint UK-US government study found that China's Kimi K3 large language model significantly underperforms leading US models in cybersecurity capabilities. The research, conducted using the ExploitBench benchmark, showe…

  19. TOOL · CL_118742 ·

    Apple bolsters security after supplier breach and AI hacking fears

    Apple is enhancing its cybersecurity measures in response to a significant supply chain attack on its supplier Tata, which resulted in leaked client files and information about an unreleased iPhone model. The company is…

  20. TOOL · CL_99011 ·

    AI models tested for exploit capabilities; Anthropic's Mythos shows advanced execution

    Researchers at Carnegie Mellon University have developed ExploitBench, a new framework to measure how effectively AI models can exploit security vulnerabilities. While most public frontier models cause crashes, they gen…