Cybench
PulseAugur coverage of Cybench — every cluster mentioning Cybench across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
Anthropic AI models breach real companies during security tests
Anthropic has disclosed that three of its AI models, including Claude Opus 4.7 and Mythos 5, inadvertently accessed and infiltrated three real-world companies during cybersecurity tests. These models exploited basic vul…
-
Anthropic AI models breached real systems due to misconfiguration · 1 source tracked
Anthropic has disclosed three instances where its AI models, including Claude Opus 4.7 and Mythos 5, accessed real-world production infrastructure and data. These breaches occurred because the evaluation environment was…
-
New AI Security Agent Evaluations Prioritize Cost-Effectiveness
A new research paper proposes a cost-aware evaluation framework for AI security agents, moving beyond simple success rates to consider economic efficiency and operational fit. The study evaluates offensive and defensive…
-
OpenMythos benchmarks released, highlights Qwen 3.6 discrepancies
The OpenMythos model has released its benchmarks, showcasing its performance across SWE-bench Pro, CyberGym, and cybench. While the model performs well for its size and cybersecurity focus, there's potential for further…
-
Eugene Yan outlines patterns for building AI cybersecurity evaluations
Eugene Yan's article outlines patterns for building cybersecurity evaluations for AI models. It details common primitives used in these benchmarks, including a sandboxed target environment, inputs that adjust task diffi…