PulseAugur
中
实时 17:42:58
English(EN) LLM agents patch security bugs, pass all tests, but still leave the vulnerability open [R]

LLM 代理难以修复安全漏洞,留下未修复的漏洞

开发了一个新的基准 CVE-Bench,用于评估 LLM 代理修复 Python 项目中安全漏洞的能力。在 18 个项目和 20 个真实 CVE 中,表现最好的模型在完全修复漏洞方面的成功率仅为 50%。值得注意的是,即使模型似乎修复了错误并通过了回归测试,漏洞通常仍然存在,这凸显了一种危险的故障模式,即在没有隐藏的安全测试的情况下,修复与正确修复无法区分。 AI

影响 LLM 代理在可靠修复安全漏洞方面显示出显著的局限性,表明在安全关键型应用程序中部署之前需要进行更严格的测试和开发。

排序理由 该集群描述了一个新的基准和对 LLM 代理在特定任务上的评估,这属于研究范畴。[lever_c_demoted from research: ic=1 ai=1.0]

在 r/MachineLearning 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

LLM 代理难以修复安全漏洞,留下未修复的漏洞

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群描述了一个新的基准和对 LLM 代理在特定任务上的评估,这属于研究范畴。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, safety, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
128 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. r/MachineLearning TIER_1 English(EN) · /u/Fickle-Box1433 ·

    大型语言模型代理修复安全漏洞,通过所有测试,但仍留下潜在风险 [R]

    <table> <tr><td> <a href="https://www.reddit.com/r/MachineLearning/comments/1tukvjt/llm_agents_patch_security_bugs_pass_all_tests_but/"> <img alt="LLM agents patch security bugs, pass all tests, but still leave the vulnerability open [R]" src="https://preview.redd.it/g29hj3ndzt4h…