PulseAugur
中
实时 16:34:32
English(EN) ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization

新研究探索优化代理工具链以提升AI性能 · 追踪9个来源

多篇研究论文探讨了代理工具链的优化和演进,代理工具链是决定语言模型如何与其环境交互的关键组成部分。研究引入了GUI-HARVEST、ActiveSaddler、MILO和DynaHarness等新框架,专注于自我改进、自动化课程学习和动态物理治理。这些方法旨在通过基于执行反馈和证据驱动的演进来优化提示、工具接口和控制逻辑,从而提高代理在编码、通用代理任务和机器人技术等各种任务中的性能。研究结果表明,代理系统的有效性高度依赖于模型与其工具链之间的相互作用,量身定制的工具链优化可带来显著的性能提升。 AI

影响 工具链优化方面的进展可以显著提高AI代理在各种应用中的可靠性和性能。

排序理由 多篇论文发表在arXiv上,详细介绍了代理工具链优化技术的新研究。

在 arXiv cs.MA (Multiagent) 阅读 →

AI 生成摘要 · Google Gemini · 来自 20 个来源。 我们如何撰写摘要 →

新研究探索优化代理工具链以提升AI性能 · 追踪9个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
多篇论文发表在arXiv上,详细介绍了代理工具链优化技术的新研究。
Source corroboration
20 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
5 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [20]

  1. arXiv cs.AI TIER_1 English(EN) · Yixuan Li, Yiyun Zhou, Yao Long Teng, Fuchao Yang, Yanchen Deng, Zhiyi Lyu, Xuyu Dong, Feng Chen, Bo An ·

    寻找合适的匹配:Agent任务中的模型-工具交互

    arXiv:2610.00917v1 Announce Type: new Abstract: Choosing an agent system means choosing both a language model and the harness through which it acts. We ask whether a strong model, harness, or pairing stays strong when the setting changes. We evaluate 66 configurations: four confi…

  2. arXiv cs.AI TIER_1 English(EN) · Geyi Yang, Zikun Qu, Xiang Li, Zhiyong Wang, Min Zhang, Shipei Zeng, Zhongxiang Dai ·

    GUI-HARVEST:通过证据驱动的 Harness 演进实现自改进 GUI 代理

    arXiv:2610.00948v1 Announce Type: cross Abstract: The executable harness surrounding a GUI model determines how observations are assembled, actions are executed, and verification, recovery, and termination are controlled. Compared with harness optimization for non-GUI agents, aut…

  3. arXiv cs.AI TIER_1 English(EN) · Sungho Park, Wonjoong Kim, Jue Zhang, Wook-Shin Han, Pengfei Gao, Chanyoung Park, Yongqiang Yao, Rao Fu, Elsie Nallipogu, Qingwei Lin, Victor R\"uhle ·

    ActiveSaddler:代理马具优化的自动化课程学习

    arXiv:2610.00906v1 Announce Type: new Abstract: Automated harness optimization can substantially improve LLM agents by iteratively updating their prompts, tool interfaces, and control logic from execution feedback. However, existing methods primarily optimize how the harness is u…

  4. arXiv cs.AI TIER_1 English(EN) · Chen Lu, Ke Xue, Siyuan Xu, Mingxuan Yuan, Chao Qian ·

    自我演化算法设计代理:通过种群策选策略优化逃离上下文内演化停滞

    arXiv:2609.38757v1 Announce Type: new Abstract: Large language models are increasingly participating in complex real-world tasks in the form of algorithm-design agents, designing and refining algorithms. Many successful algorithm-design agents adopt pure in-context evolutionary f…

  5. arXiv cs.AI TIER_1 English(EN) · Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Bl\"obaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc ·

    MILO:通过协调多智能体进化实现自动化线束发现

    arXiv:2609.38349v1 Announce Type: cross Abstract: Modern agentic systems combine an AI model with a harness that controls execution and environmental interactions. Harness design strongly affects long-horizon performance, yet its combinatorial search space demands substantial hum…

  6. arXiv cs.AI TIER_1 English(EN) · Zhixuan Tan, Pengjie Gu, Zhao Li, Yihan Hu, Xu He, Dong Li, Jianye Hao ·

    EngramBench: 一个基于能力演进的基准测试工具

    arXiv:2609.39284v1 Announce Type: cross Abstract: While large language models have achieved remarkable success in isolated code generation, authentic software engineering requires sustained reasoning, complex state management, and continuous cross-domain abstraction. However, cur…

  7. arXiv cs.AI TIER_1 English(EN) · Zhijie Wei, Ferris Tan, Jinghui Wang ·

    规模与选择:是什么让自动线束进化适用于视觉界面机器人代理

    arXiv:2609.39304v1 Announce Type: cross Abstract: When an off-the-shelf coding agent is used directly as a robot policy, observing a browser-based 3D interface through screenshots and acting by posing a virtual target gripper through a few tools, the agent's harness, its prompts,…

  8. arXiv cs.AI TIER_1 English(EN) · Minki Kang, Ryo Hachiuma, Shaokun Zhang, Subhashree Radhakrishnan, Yonggan Fu, Jindong Jiang, Mingjie Liu, Ehsan Hosseini-Asl, Yi Dong, Yu-Chiang Frank Wang, Byung-Kwan Lee ·

    Mid-Harness:在模型和Harness之间扩展用于终端代理的操作

    arXiv:2609.39982v1 Announce Type: cross Abstract: Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution. A poor command (e.g., wrong package install) can change the environment in ways that hin…

  9. arXiv cs.AI TIER_1 English(EN) · Haoyuan Deng, Jiebin Liu, Tengxiao Zhang, Langning Yan, Hongye Cao, Ziwei Wang ·

    DynaHarness:一种用于自进化机器人代理的动态物理线束

    arXiv:2609.40306v1 Announce Type: cross Abstract: Pretrained robot policies provide useful action priors, but long-horizon manipulation still requires coordination between semantic reasoning and physical execution. Semantic reasoning operates at a coarser timescale than physical …

  10. arXiv cs.AI TIER_1 English(EN) · Zhening Li, Joshua Liu, Mateja Vukelic, Nicole Shen, Supriya Lall, Amitayush Thakur, Alex Zhang, Omar Khattab, Jonathan Light, Armando Solar-Lezama ·

    Harness 语言:一个具有最大表现力的极简代理框架

    arXiv:2609.26891v2 Announce Type: replace Abstract: Modern language-model agents are built around the agent loop: the LLM is placed in an environment exposing a set of tools, and the LLM has full control over the workflow by alternating between tool calls and observing their outp…

  11. arXiv cs.AI TIER_1 English(EN) · Qiankai Xu ·

    Self-Evolving Harness on Multiple Tasks with the Agent as Its Own Optimizer

    arXiv:2609.38372v1 Announce Type: new Abstract: A harness is the code around a language-model agent that organizes prompts, calls tools, manages context, and controls execution. As models grow stronger, recent work has begun to let agents improve their own harnesses, a line of wo…

  12. arXiv cs.AI TIER_1 English(EN) · Jingbo Yang, Kwei-Herng Lai, Xiaowen Wang, Yaar Harari, Evgeniy Gabrilovich, Shiyu Chang ·

    从研究中学习:迈向终身智能体演化

    arXiv:2609.40169v1 Announce Type: new Abstract: Language agents are expected to solve increasingly complex tasks, creating a growing need for continual improvement. One promising approach is to evolve the agent harness, the software that governs tool use, memory management, and t…

  13. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Victor Rühle ·

    ActiveSaddler:用于智能体约束优化的自动化课程学习

    Automated harness optimization can substantially improve LLM agents by iteratively updating their prompts, tool interfaces, and control logic from execution feedback. However, existing methods primarily optimize how the harness is updated while largely fixing which training scena…

  14. Hugging Face Daily Papers TIER_1 English(EN) ·

    ActiveSaddler:用于智能体约束优化的自动化课程学习

    Automated harness optimization can substantially improve LLM agents by iteratively updating their prompts, tool interfaces, and control logic from execution feedback. However, existing methods primarily optimize how the harness is updated while largely fixing which training scena…

  15. Hugging Face Daily Papers TIER_1 English(EN) ·

    从研究中学习:迈向终身智能体演化

    Language agents are expected to solve increasingly complex tasks, creating a growing need for continual improvement. One promising approach is to evolve the agent harness, the software that governs tool use, memory management, and task execution, while keeping the underlying lang…

  16. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Byung-Kwan Lee ·

    Mid-Harness:在模型和Harness之间扩展用于终端代理的操作

    Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution. A poor command (e.g., wrong package install) can change the environment in ways that hinder subsequent progress, even when the model could…

  17. arXiv cs.AI TIER_1 English(EN) · Haoyu Dong, Yuhang Zhou, Zihao Lin, Yifan Wu, Bo Peng, Mingyi Wang, Xiangjun Fan, Lizhu Zhang, Zhuokai Zhao ·

    用于智能体工具链优化的自改进分支混合模型

    arXiv:2609.37834v1 Announce Type: new Abstract: Harness optimization provides a practical setting for recursive self-improvement (RSI), where agent-generated modifications inform subsequent changes through execution feedback. Recent work such as Meta-Harness implements this proce…

  18. arXiv cs.NE (Neural & Evolutionary) TIER_1 English(EN) · Chao Qian ·

    自主进化算法设计代理:通过种群精选策略优化逃离上下文内进化停滞

    Large language models are increasingly participating in complex real-world tasks in the form of algorithm-design agents, designing and refining algorithms. Many successful algorithm-design agents adopt pure in-context evolutionary frameworks, but they may quickly plateau in domai…

  19. Hugging Face Daily Papers TIER_1 English(EN) ·

    Mid-Harness:在模型和Harness之间扩展用于终端代理的操作

    Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution. A poor command (e.g., wrong package install) can change the environment in ways that hinder subsequent progress, even when the model could…

  20. Hugging Face Daily Papers TIER_1 English(EN) ·

    MILO:通过协调多智能体进化实现自动化线束发现

    Modern agentic systems combine an AI model with a harness that controls execution and environmental interactions. Harness design strongly affects long-horizon performance, yet its combinatorial search space demands substantial human effort that must be repeated as models change. …