研究人员开发了一种新的强化学习方法来训练人工智能代理进行元推理,使其能够决定是进行快速的反应式决策还是缓慢的审议式规划。该方法使用一种元推理策略,预测反应式策略何时可能表现不佳,从而发出信号表明需要通过规划进行更多计算。在运动规划和导航环境中的实验表明,该系统可以有效地学习何时进行规划,而不是依赖于反应式策略,并且它能够适应以提高反应式策略的性能。 AI
影响 这项研究可能带来更高效、更适应性强的人工智能系统,能够优化决策的计算资源。
排序理由 该集群包含一篇详细介绍新方法的学术论文。 [lever_c_demoted from research: ic=1 ai=1.0]
- artificial intelligence
- arXiv
- imitation learning
- Meta-Reasoning: Monitoring and Control of Thinking and Reasoning
- motion planning
- navigation environments
- reinforcement learning
- When to Plan: Learning to Select Between Reactive Control and Deliberative Planning
AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →