PulseAugur
实时 14:33:08
English(EN) When to Plan: Learning to Select Between Reactive Control and Deliberative Planning

新的强化学习方法教会 AI 代理何时规划或反应

研究人员开发了一种新的强化学习方法来训练人工智能代理进行元推理,使其能够决定是进行快速的反应式决策还是缓慢的审议式规划。该方法使用一种元推理策略,预测反应式策略何时可能表现不佳,从而发出信号表明需要通过规划进行更多计算。在运动规划和导航环境中的实验表明,该系统可以有效地学习何时进行规划,而不是依赖于反应式策略,并且它能够适应以提高反应式策略的性能。 AI

影响 这项研究可能带来更高效、更适应性强的人工智能系统,能够优化决策的计算资源。

排序理由 该集群包含一篇详细介绍新方法的学术论文。 [lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的强化学习方法教会 AI 代理何时规划或反应

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Adam Labiosa, Josiah P. Hanna ·

    何时规划:学习在反应式控制与审慎规划之间进行选择

    arXiv:2607.16421v1 Announce Type: new Abstract: It has long been recognized that humans have the ability to switch between fast, reactive decision-making and slower, deliberative planning. In this paper, we study the question of how to learn this ability, known as meta-reasoning,…