PulseAugur
实时 10:55:08
English(EN) TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents

新的TRACE方法增强了AI智能体在长时程任务中的工具使用 · 跟踪2个来源

研究人员开发了TRACE,一种用于提高多回合AI智能体在复杂、长时程任务中性能的新方法。该技术通过从参考模型的对数概率中推导出每个动作的奖励,而不是仅仅依赖稀疏的结果奖励,来解决信用分配的挑战。TRACE显著提升了Qwen3-4B和Qwen3-30B-A3B等模型在BrowseComp-Plus等基准测试中的工具使用能力,从而加快了收敛速度并改善了学习曲线。 AI

影响 增强了AI智能体在复杂、多回合任务中的能力,可能加速需要长时程推理的领域的进展。

排序理由 该集群包含一篇详细介绍AI智能体新方法的论文。

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

新的TRACE方法增强了AI智能体在长时程任务中的工具使用 · 跟踪2个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群包含一篇详细介绍AI智能体新方法的论文。
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
63 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准

报道来源 [2]

  1. arXiv cs.LG TIER_1 English(EN) · Leitian Tao, Baolin Peng, Wenlin Yao, Tao Ge, Hao Cheng, Mike Hang Wang, Jianfeng Gao, Sharon Li ·

    TRACE:通过信用估计为长时域智能体分配回合级奖励

    arXiv:2607.13988v1 Announce Type: new Abstract: Multi-turn agents solve complex tasks through extended sequences of tool interactions before producing a final answer, making credit assignment a fundamental challenge during post-training. Outcome rewards provide reliable supervisi…

  2. arXiv cs.LG TIER_1 English(EN) · Sharon Li ·

    TRACE:通过信用估计实现长时域智能体的回合级奖励分配

    Multi-turn agents solve complex tasks through extended sequences of tool interactions before producing a final answer, making credit assignment a fundamental challenge during post-training. Outcome rewards provide reliable supervision for short-horizon reasoning, but become spars…