Researchers have introduced VICBench, a new benchmark designed to evaluate code vulnerability detection tools. The benchmark comprises 100 verified vulnerability-inducing commits (VICs) across 88 projects in Python, Java, and C++, covering 48 Common Weakness Enumeration (CWE) types. VICBench features complex, real-world vulnerability fixes and their corresponding VICs, which are significantly larger than those in previous datasets. Evaluations using VICBench indicate that current state-of-the-art algorithms like V-SZZ and LLM4SZZ achieve only moderate F1 scores, highlighting the need for further development in automated vulnerability detection. AI
IMPACT This benchmark will enable more robust evaluation of AI-driven code vulnerability detection, potentially accelerating the development of more effective security tools.
RANK_REASON The cluster describes a new academic benchmark for evaluating AI models in code vulnerability detection.
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →