PulseAugur
EN
LIVE 15:34:07
Deutsch(DE) Qwen 3.8 abliterated in CyBench: Autonome Cyberfähigkeiten auf dem Prüfstand Blackfrost löst 18 von 39 CTF-Aufgaben autonom, scheitert aber an Reasoning-Loops u

Qwen 3.8 struggles in autonomous cyber benchmark; Blackfrost completes 18/39 challenges

The Qwen 3.8 model was evaluated on the CyBench benchmark, which tests autonomous cyber capabilities. In the evaluation, a system named Blackfrost successfully completed 18 out of 39 Capture The Flag (CTF) challenges autonomously. However, Blackfrost encountered issues with reasoning loops and time limits, indicating limitations in its autonomous performance. AI

IMPACT Evaluates the autonomous cyber capabilities of AI models, highlighting limitations in reasoning and time management.

RANK_REASON The cluster reports on the results of a benchmark evaluation of an AI model's autonomous capabilities. [lever_c_demoted from research: ic=1 ai=1.0]

Read on Mastodon — mastodon.social →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Qwen 3.8 struggles in autonomous cyber benchmark; Blackfrost completes 18/39 challenges

COVERAGE [1]

  1. Mastodon — mastodon.social TIER_1 Deutsch(DE) · aisyndicate ·

    Qwen 3.8 obliterated in CyBench: Autonomous Cyber Capabilities on the Test Stand Blackfrost autonomously solves 18 out of 39 CTF tasks, but fails at Reasoning Loops

    Qwen 3.8 abliterated in CyBench: Autonome Cyberfähigkeiten auf dem Prüfstand Blackfrost löst 18 von 39 CTF-Aufgaben autonom, scheitert aber an Reasoning-Loops und Zeitlimits. Der Test trennt bloße Willigkeit von tatsächlicher cyberautonomer Leistungsfähigkeit. https:// aisyndicat…