PulseAugur
中
实时 22:39:10
English(EN) New benchmark on LMs fixing bugs before users run into them

新的 SWE-sweep 基准测试评估大型语言模型主动修复 Bug 的能力

来自 Meta、Stanford、Harvard 和 UW 的研究人员开发了 SWE-sweep,这是一个新的基准测试,旨在评估大型语言模型在影响用户之前主动识别和修复大型代码库中 Bug 的能力。该基准测试使用真实世界的 Bug,并根据模型在给定代码库中查找和解决这些问题的成功程度对其进行评分。早期结果表明,虽然一些模型表现不佳,但 Luna xhigh 显示出成本效益,研究人员正在寻求推荐其他开源模型以纳入未来的更新。 AI

影响 该基准测试有望推动能够主动检测错误的更强大的 AI 编码助手的发展。

排序理由 学术研究人员发布了新的基准测试和论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的 SWE-sweep 基准测试评估大型语言模型主动修复 Bug 的能力

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
学术研究人员发布了新的基准测试和论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
1 days old
Coverage has settled into its steady-state source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/klieret ·

    新基准测试:LMs 在用户发现前修复 Bug

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1wvxph8/new_benchmark_on_lms_fixing_bugs_before_users_run/"> <img alt="New benchmark on LMs fixing bugs before users run into them" src="https://preview.redd.it/irffy7x5v2th1.png?width=140&amp;height=139&amp;a…