Researchers have introduced VICBench, a new benchmark designed to evaluate code vulnerability detection tools. This benchmark comprises 100 verified vulnerability-inducing commits (VICs) across 88 projects in Python, Java, and C++, covering 48 Common Weakness Enumeration (CWE) types. VICBench features complex, real-world fixes and their corresponding VICs, which are significantly larger than those found in prior work. Evaluations using VICBench indicate that current state-of-the-art algorithms like V-SZZ and LLM4SZZ achieve only moderate F1 scores, highlighting the continued need for manual effort in vulnerability detection. AI
IMPACT This benchmark will enable more robust evaluation of AI-driven code vulnerability detection, potentially accelerating the development of more effective security tools.
RANK_REASON The item describes a new academic paper introducing a benchmark for evaluating AI models in code vulnerability detection. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →