PulseAugur
实时 21:31:44
English(EN) Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails

AI模型通过协同进化的约束和定向纠正进行学习 · 跟踪2个来源

研究人员开发了一种新颖的方法,通过协同进化其“约束”(系统提示、工具集和脚手架)和权重来提高较小AI模型在特定任务上的性能。他们发现,直接模仿专家轨迹会破坏模型的原生规划风格,从而降低性能。为解决此问题,他们引入了一个策略内专家纠正流程,仅识别并重写弱模型自身回放中的失败回合,保留其规划风格,并为领域特定的企业任务实现经济高效的协同进化。 AI

影响 这项研究通过优化AI模型与工具和提示的交互方式,为提高其在专业任务上的性能提供了一种更经济高效的方法。

排序理由 该集群包含一篇详细介绍AI模型训练新方法的学术论文。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

AI模型通过协同进化的约束和定向纠正进行学习 · 跟踪2个来源

本文如何被排名

Signal score
1 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群包含一篇详细介绍AI模型训练新方法的学术论文。
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
2 days old
Coverage has settled into its steady-state source set.

完整方法见我们的编辑标准

报道来源 [2]

  1. arXiv cs.AI TIER_1 English(EN) · Zhou Yu, Bin Bi, Shiva Kumar Pentyala, Shubham Mehrotra, Sougata Chaudhuri, Shilpa Bhagavath, Zeyuan Chen, Ran Xu, Phil Mui, James Zhu, Sitaram Asur ·

    协同进化中的 Harness 和模型:On-Policy Correction 助力弱模型在模仿失效时迎头赶上

    arXiv:2609.09134v1 Announce Type: new Abstract: Agent harnesses (the system prompt, tool set, execution hooks, and context-management scaffolding around a model) are a critical determinant of agentic task success. Automated harness evolution can enable smaller models to perform w…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    协同进化中的 Harness 和模型:On-Policy Correction 帮助较弱模型在模仿失败的地方迎头赶上

    Combining harness evolution with localized expert correction improves weaker models without disrupting their native planning style.