Researchers have developed HoF-Bench, a new benchmark designed to evaluate AI models' ability to find real-world software vulnerabilities without relying on frontier models. The benchmark consists of 95 AI-discovered CVEs from mature open-source projects like OpenSSL and curl. In testing, a deliberately minimal LLM-based analyzer successfully rediscovered 68% of these vulnerabilities under strict conditions, while no frontier models were used for detection in the study. HoF-Bench aims to provide a standardized method for comparing vulnerability scanners and assessing their reliability, particularly noting that vulnerabilities in C programming language code proved more challenging for the tested models. AI
IMPACT This benchmark could accelerate the development of more reliable AI-powered security tools for identifying software vulnerabilities.
RANK_REASON The cluster describes a new benchmark for evaluating AI models on a specific research task (finding software vulnerabilities), based on an academic paper. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →