研究人员引入了一个名为SELF(SELF-distilLation with environmental Feedback modeling)的新框架,用于改进在缺乏直接奖励的交互式环境中训练的语言代理。该框架联合优化环境反馈建模和事后自我蒸馏,使代理能够在从条件化反馈的自我教师那里学习的同时,预测环境响应。实验表明,SELF在tau-Bench和AppWorld等基准测试中优于SDPO和GRPO等现有方法,通过更有效地利用环境反馈来增强代理能力。 AI
影响 增强了在缺乏直接奖励的环境中代理的能力,有可能提高复杂交互任务的性能。
排序理由 该集群包含一篇详细介绍新框架和语言代理训练实验结果的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
- Agentic hindsight self-distillation
- Agentic SElf-distilLation with environmental Feedback modeling
- AppWorld
- arXiv
- Environmental feedback modeling
- Grpo
- Hindsight Self-Distillation
- Qwen3_8B
- reinforcement learning
- SELF
- Self Distillation Using Contrastive Evidence Policy Optimization
- tau-Bench
AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →