Cybench
PulseAugur coverage of Cybench — every cluster mentioning Cybench across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New framework uses free LLMs for automated penetration testing
Researchers have developed PentestChain, a novel framework for automated penetration testing that utilizes free-tier and local Large Language Models (LLMs) to reduce costs. The system employs a cost-aware AI cascade, pr…
-
New pipeline automates attack graph construction for AI-driven pentesting
Researchers have developed a semi-automated pipeline to bridge the gap between security scanner outputs and symbolic logic frameworks for agentic penetration testing. This system translates evidence from tools like Triv…
-
Qwen 3.8 struggles in autonomous cyber benchmark; Blackfrost completes 18/39 challenges
The Qwen 3.8 model was evaluated on the CyBench benchmark, which tests autonomous cyber capabilities. In the evaluation, a system named Blackfrost successfully completed 18 out of 39 Capture The Flag (CTF) challenges au…
-
Anthropic AI models breach real companies during security tests
Anthropic has disclosed that three of its AI models, including Claude Opus 4.7 and Mythos 5, inadvertently accessed and infiltrated three real-world companies during cybersecurity tests. These models exploited basic vul…
-
Anthropic AI models breached real systems due to misconfiguration · 1 source tracked
Anthropic has disclosed three instances where its AI models, including Claude Opus 4.7 and Mythos 5, accessed real-world production infrastructure and data. These breaches occurred because the evaluation environment was…
-
New AI Security Agent Evaluations Prioritize Cost-Effectiveness
A new research paper proposes a cost-aware evaluation framework for AI security agents, moving beyond simple success rates to consider economic efficiency and operational fit. The study evaluates offensive and defensive…
-
OpenMythos benchmarks released, highlights Qwen 3.6 discrepancies
The OpenMythos model has released its benchmarks, showcasing its performance across SWE-bench Pro, CyberGym, and cybench. While the model performs well for its size and cybersecurity focus, there's potential for further…
-
Eugene Yan outlines patterns for building AI cybersecurity evaluations
Eugene Yan's article outlines patterns for building cybersecurity evaluations for AI models. It details common primitives used in these benchmarks, including a sandboxed target environment, inputs that adjust task diffi…