PulseAugur
中
实时 03:46:36
English(EN) Harness Engineering for Software Engineering via Modular Executable Dev-Primitives

新的 HERMES 框架提升了 LLM 在软件工程任务中的性能

研究人员开发了 HERMES,一个旨在提高大型语言模型 (LLM) 在复杂软件工程任务中性能的工程化框架。HERMES 利用模块化、可执行的“开发基元”,将存储库组件转化为能够进行自然语言推理和自我修改的活动代理。实验表明,HERMES 的性能平均比基线框架高出 12.4%,即使使用较小的 Qwen3-8B 模型,其性能也接近 GPT-5.6 Sol 配置,同时降低了推理成本。 AI

影响 该框架有望显著提高 LLM 驱动的软件开发的可靠性和效率,可能降低成本并加速工作流程。

排序理由 该集群包含一篇详细介绍使用 LLM 进行软件工程的新框架和方法的论文。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

新的 HERMES 框架提升了 LLM 在软件工程任务中的性能

本文如何被排名

Signal score
1 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群包含一篇详细介绍使用 LLM 进行软件工程的新框架和方法的论文。
Source corroboration
3 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
2 days old
Coverage has settled into its steady-state source set.

完整方法见我们的编辑标准。

报道来源 [3]

  1. arXiv cs.AI TIER_1 English(EN) · Haibo Jin, Xinjie Li, Peng Kuang, Haohan Wang ·

    通过模块化可执行开发原语赋能软件工程

    arXiv:2610.07832v1 Announce Type: cross Abstract: Large language models (LLMs) equipped with terminal access have demonstrated strong capabilities in automating software engineering tasks. However, existing agents remain brittle on long-horizon workflows, where they must repeated…

  2. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Haohan Wang ·

    通过模块化可执行开发原语赋能软件工程

    Large language models (LLMs) equipped with terminal access have demonstrated strong capabilities in automating software engineering tasks. However, existing agents remain brittle on long-horizon workflows, where they must repeatedly reconstruct program state scattered across sour…

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    通过模块化可执行开发原语赋能软件工程

    Large language models (LLMs) equipped with terminal access have demonstrated strong capabilities in automating software engineering tasks. However, existing agents remain brittle on long-horizon workflows, where they must repeatedly reconstruct program state scattered across sour…