English(EN)ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization
新研究探索优化代理工具链以提升AI性能 · 追踪9个来源
作者PulseAugur 编辑部·[20 个来源]·
多篇研究论文探讨了代理工具链的优化和演进,代理工具链是决定语言模型如何与其环境交互的关键组成部分。研究引入了GUI-HARVEST、ActiveSaddler、MILO和DynaHarness等新框架,专注于自我改进、自动化课程学习和动态物理治理。这些方法旨在通过基于执行反馈和证据驱动的演进来优化提示、工具接口和控制逻辑,从而提高代理在编码、通用代理任务和机器人技术等各种任务中的性能。研究结果表明,代理系统的有效性高度依赖于模型与其工具链之间的相互作用,量身定制的工具链优化可带来显著的性能提升。
AI
arXiv:2610.00917v1 Announce Type: new Abstract: Choosing an agent system means choosing both a language model and the harness through which it acts. We ask whether a strong model, harness, or pairing stays strong when the setting changes. We evaluate 66 configurations: four confi…
arXiv:2610.00948v1 Announce Type: cross Abstract: The executable harness surrounding a GUI model determines how observations are assembled, actions are executed, and verification, recovery, and termination are controlled. Compared with harness optimization for non-GUI agents, aut…
arXiv cs.AI
TIER_1English(EN)·Sungho Park, Wonjoong Kim, Jue Zhang, Wook-Shin Han, Pengfei Gao, Chanyoung Park, Yongqiang Yao, Rao Fu, Elsie Nallipogu, Qingwei Lin, Victor R\"uhle·
arXiv:2610.00906v1 Announce Type: new Abstract: Automated harness optimization can substantially improve LLM agents by iteratively updating their prompts, tool interfaces, and control logic from execution feedback. However, existing methods primarily optimize how the harness is u…
arXiv:2609.38757v1 Announce Type: new Abstract: Large language models are increasingly participating in complex real-world tasks in the form of algorithm-design agents, designing and refining algorithms. Many successful algorithm-design agents adopt pure in-context evolutionary f…
arXiv:2609.38349v1 Announce Type: cross Abstract: Modern agentic systems combine an AI model with a harness that controls execution and environmental interactions. Harness design strongly affects long-horizon performance, yet its combinatorial search space demands substantial hum…
arXiv:2609.39284v1 Announce Type: cross Abstract: While large language models have achieved remarkable success in isolated code generation, authentic software engineering requires sustained reasoning, complex state management, and continuous cross-domain abstraction. However, cur…
arXiv:2609.39304v1 Announce Type: cross Abstract: When an off-the-shelf coding agent is used directly as a robot policy, observing a browser-based 3D interface through screenshots and acting by posing a virtual target gripper through a few tools, the agent's harness, its prompts,…
arXiv:2609.39982v1 Announce Type: cross Abstract: Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution. A poor command (e.g., wrong package install) can change the environment in ways that hin…
arXiv:2609.40306v1 Announce Type: cross Abstract: Pretrained robot policies provide useful action priors, but long-horizon manipulation still requires coordination between semantic reasoning and physical execution. Semantic reasoning operates at a coarser timescale than physical …
arXiv:2609.26891v2 Announce Type: replace Abstract: Modern language-model agents are built around the agent loop: the LLM is placed in an environment exposing a set of tools, and the LLM has full control over the workflow by alternating between tool calls and observing their outp…
arXiv:2609.38372v1 Announce Type: new Abstract: A harness is the code around a language-model agent that organizes prompts, calls tools, manages context, and controls execution. As models grow stronger, recent work has begun to let agents improve their own harnesses, a line of wo…
arXiv:2609.40169v1 Announce Type: new Abstract: Language agents are expected to solve increasingly complex tasks, creating a growing need for continual improvement. One promising approach is to evolve the agent harness, the software that governs tool use, memory management, and t…
Automated harness optimization can substantially improve LLM agents by iteratively updating their prompts, tool interfaces, and control logic from execution feedback. However, existing methods primarily optimize how the harness is updated while largely fixing which training scena…
Automated harness optimization can substantially improve LLM agents by iteratively updating their prompts, tool interfaces, and control logic from execution feedback. However, existing methods primarily optimize how the harness is updated while largely fixing which training scena…
Language agents are expected to solve increasingly complex tasks, creating a growing need for continual improvement. One promising approach is to evolve the agent harness, the software that governs tool use, memory management, and task execution, while keeping the underlying lang…
Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution. A poor command (e.g., wrong package install) can change the environment in ways that hinder subsequent progress, even when the model could…
arXiv:2609.37834v1 Announce Type: new Abstract: Harness optimization provides a practical setting for recursive self-improvement (RSI), where agent-generated modifications inform subsequent changes through execution feedback. Recent work such as Meta-Harness implements this proce…
Large language models are increasingly participating in complex real-world tasks in the form of algorithm-design agents, designing and refining algorithms. Many successful algorithm-design agents adopt pure in-context evolutionary frameworks, but they may quickly plateau in domai…
Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution. A poor command (e.g., wrong package install) can change the environment in ways that hinder subsequent progress, even when the model could…
Modern agentic systems combine an AI model with a harness that controls execution and environmental interactions. Harness design strongly affects long-horizon performance, yet its combinatorial search space demands substantial human effort that must be repeated as models change. …