PulseAugur
中
实时 07:00:00
English(EN) Trustworthy Runtime Error Healing in Real-World Repositories: A Benchmark and Guardrail

用于 LLM 驱动的软件运行时错误修复的新基准与防护栏

研究人员开发了 HealBench,这是一个旨在评估大型语言模型 (LLM) 在真实软件代码库中自动修复运行时错误的有效性和安全性新基准。该系统包括 HealGuard,一种使用静态和动态污点分析的安全机制,以确保 LLM 生成的代码不会损害受保护的操作。在评估中,表现最佳的 LLM 设置在 38.11% 的实例中成功恢复了执行,并在 28.68% 的情况下通过了目标测试,尽管 HealGuard 标记了其中很大一部分可能不安全。 AI

影响 这项研究提高了 LLM 在自动代码修复方面的安全性与可靠性,有望实现更强大的软件开发工具。

排序理由 该项目是一篇研究论文,介绍了一个用于基于 LLM 的软件错误修复的新基准和安全机制。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

用于 LLM 驱动的软件运行时错误修复的新基准与防护栏

本文如何被排名

Signal score
25 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该项目是一篇研究论文,介绍了一个用于基于 LLM 的软件错误修复的新基准和安全机制。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, safety, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Gou Tan, Pengfei Chen, Zhensu Sun, Jieke Shi, Junkai Chen, Ting Zhang, Weifeng Sun, Junda He, Shuai Liang, Chuanfu Zhang, Lwin Khin Shar, David Lo ·

    真实代码库中可信赖的运行时错误修复:基准测试与护栏

    arXiv:2609.39086v1 Announce Type: cross Abstract: Runtime error healing lets a crashed program continue by generating code that repairs its live runtime state. Recent work shows that LLMs can generate such healing code, but it is evaluated only on small competition programs, and …