Researchers have introduced the Vulnerability Localization Benchmark (VLoc Bench), a new dataset designed to evaluate the ability of AI agents to identify specific code locations associated with security vulnerabilities within software repositories. The benchmark, comprising 500 real-world vulnerabilities across 290 repositories and 147 Common Weakness Enumeration (CWE) categories, tests agents on their capacity to pinpoint affected files. Initial evaluations of 27 language models and four static-analysis tools revealed that repository-scale vulnerability localization remains a significant challenge, with the top-performing system achieving only a 0.229 File F1 score. AI
IMPACT Establishes a new evaluation standard for AI security agents, highlighting current limitations in code-level vulnerability identification.
RANK_REASON The cluster describes a new academic benchmark for evaluating AI agents on a specific task, published on arXiv. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Common Weakness Enumeration
- DagsHub
- File F1
- Hugging Face
- VLoc Bench
- Vulnerability Localization Benchmark
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →