PulseAugur
中
实时 06:54:32
English(EN) SWE-Together: Evaluating Coding Agents in Interactive User Sessions

新的基准测试评估了交互式、多轮会话中的AI编码代理

引入了两个新的基准测试SWE-Together和SWE-Interact,用于在更真实、交互式和多轮的用户会话中评估编码代理。与一次性提供完整任务描述的静态基准测试不同,这些新框架模拟用户交互,逐步揭示需求并提供反馈。实验表明,在单轮任务上的强劲表现并不总是能转化为多轮场景,像Opus 4.8和GPT 5.5这样的顶级模型仍然存在忘记需求或犯技术性错误等问题。 AI

影响 这些基准测试将推动更强大的AI编码助手的发展,使其能够处理复杂的交互式软件工程任务。

排序理由 两篇学术论文介绍了用于评估AI编码代理的新基准测试。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 5 个来源。 我们如何撰写摘要 →

新的基准测试评估了交互式、多轮会话中的AI编码代理

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
两篇学术论文介绍了用于评估AI编码代理的新基准测试。
Source corroboration
5 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
93 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [5]

  1. arXiv cs.AI TIER_1 English(EN) · Yifan Wu, Zhuokai Zhao, Songlin Li, Ho Hin Lee, Jiacheng Zhu, Shirley Wu, Tianhe Yu, Serena Li, Lizhu Zhang, Xiangjun Fan, Shengzhi Li ·

    SWE-Together:在交互式用户会话中评估编码代理

    arXiv:2606.29957v1 Announce Type: cross Abstract: Most coding-agent benchmarks are static: an agent receives a complete task description up front and is judged only by its final code. Real coding assistance is interactive, with users clarifying goals, adding constraints, and corr…

  2. arXiv cs.LG TIER_1 English(EN) · Mohit Raghavendra, Anisha Gunjal, Aakash Sabharwal, Yunzhong He ·

    SWE-INTERACT:将 SWE 基准重新构想为用户驱动的长程编码会话

    arXiv:2606.30573v1 Announce Type: new Abstract: We introduce SWE-Interact, a new testbed for evaluating coding agents on multi-turn, interactive, user-driven software engineering tasks. Existing frontier SWE benchmarks typically provide complete requirements upfront and evaluate …

  3. arXiv cs.LG TIER_1 English(EN) · Yunzhong He ·

    SWE-INTERACT:将 SWE 基准重新构想为用户驱动的远程编码会话

    We introduce SWE-Interact, a new testbed for evaluating coding agents on multi-turn, interactive, user-driven software engineering tasks. Existing frontier SWE benchmarks typically provide complete requirements upfront and evaluate agents on autonomous implementation. In contrast…

  4. Hugging Face Daily Papers TIER_1 English(EN) ·

    SWE-Together:在交互式用户会话中评估编码代理

    SWE-Together is a multi-turn coding benchmark created from real user-agent interactions, featuring a reactive LLM simulator to evaluate agents based on both final correctness and interaction efficiency.

  5. Hugging Face Daily Papers TIER_1 English(EN) ·

    SWE-INTERACT:将 SWE 基准重新构想为用户驱动的长时程编码会话

    SWE-Interact presents a testbed that evaluates coding agents in realistic multi-turn, user-driven software engineering scenarios, revealing significant gaps between single-turn performance and interactive task completion.