Researchers have introduced FuzzingBrain-Bench V1, a new benchmark designed to evaluate the open-ended bug discovery capabilities of large language models (LLMs). Unlike previous benchmarks that focus on triggering predefined vulnerabilities, FuzzingBrain-Bench V1 challenges models to find as many distinct crashes as possible in open-source software within a Dockerized environment. The benchmark includes 77 challenges across 43 projects, with a focus on C, C++, and Java/JVM code. In evaluations, Claude Opus 4.8 demonstrated the strongest performance, successfully triggering crashes in 60 of the 77 challenges. AI
IMPACT This benchmark could lead to more robust LLM evaluations for software security and reliability.
RANK_REASON The cluster describes a new academic benchmark for evaluating LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
- Claude Haiku 4.5
- Claude Opus 4.8
- Claude Sonnet 4.6
- CPP
- C programming language
- Docker
- FuzzingBrain-Bench V1
- Java
- Java virtual machine
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →