PulseAugur
中
实时 06:55:47
English(EN) RustMizan: A Compilable, Contamination-Aware Benchmarking Framework for Rust Vulnerabilities

发布用于 LLM 的 Java 和 Rust 漏洞检测新基准

发布了两个新的基准测试集 JavaVulBench 和 RustMizan,用于评估大型语言模型在软件漏洞检测方面的能力。JavaVulBench 专注于 Java 方法,包含超过 1,740 个通用漏洞披露 (CVE),并提供多种真实的拆分策略用于测试。RustMizan 针对 Rust 漏洞,提供可编译的代码和一个突变框架来测试污染和鲁棒性。与之前使用小型代码片段且缺乏污染意识的数据集相比,这两个基准测试旨在提供更现实、更全面的评估。 AI

影响 这些基准测试将能够对用于代码安全的 LLM 进行更严格的评估,可能加速 AI 在软件开发安全领域的应用。

排序理由 该集群包含两篇研究论文,介绍了用于评估 AI 模型在软件漏洞检测方面的新基准测试。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

发布用于 LLM 的 Java 和 Rust 漏洞检测新基准

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群包含两篇研究论文,介绍了用于评估 AI 模型在软件漏洞检测方面的新基准测试。
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
93 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [2]

  1. arXiv cs.AI TIER_1 English(EN) · Norbert Sandor Szolnoki, Gabor Antal ·

    JavaVulBench:一个具有真实划分、统一多后端工具集和泄露感知评估模式的Java漏洞基准测试

    arXiv:2607.02825v1 Announce Type: cross Abstract: We release \textsc{JavaVulBench}, a benchmark dataset and evaluation harness for Java vulnerability detection. The dataset contains $\sim$30{,}600 Java methods spanning 1{,}740 CVEs and 700+ projects, labelled at both method and l…

  2. arXiv cs.AI TIER_1 English(EN) · Tarek Elsayed, Shiping Yang, Eunsong Koh, Sanika Goyal, Vincent Huang, Paul Ngo, Nathan Young, Mohammad Omidvar Tehrani, Alvyn Kang, Arnell Kang, Zeyu Chen, Ang\'elica Moreira, Xuan Feng, Angel X. Chang, Nick Sumner, Steven Y. Ko ·

    RustMizan: 一个可编译、可感知污染的 Rust 漏洞基准测试框架

    arXiv:2607.04729v1 Announce Type: cross Abstract: LLM agents are increasingly applied to vulnerability analysis, but existing benchmarks have not kept pace. They typically rely on small non-compilable snippets, focus on binary classification (vulnerable or not), and do not accoun…