PulseAugur
实时 09:25:36
English(EN) Generalised Bellman recurrence and three dualities in sequential decision-making

新的arXiv论文探讨基于风险的MDP和贝尔曼方程对偶

两篇新发表在arXiv上的研究论文探讨了人工智能序贯决策中的高级概念。第一篇论文介绍了ERQDP,一种在基于风险的目标下进行有限时间马尔可夫决策过程规划的新方法,它提供了一种无需枚举的方法,并提供认证解决方案或明确的剩余差距。第二篇论文深入研究了贝尔曼方程的理论基础,展示了其递推性质如何源于与状态动力学、回报分解和不确定性聚合相关的三个基本条件,统一了强化学习、控制和决策理论中的概念。 AI

影响 这些论文推进了人工智能决策的理论框架,有可能提高人工智能代理在复杂环境中的鲁棒性和效率。

排序理由 两篇发表在arXiv上的学术论文,详细介绍了人工智能决策方面的理论进展。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

新的arXiv论文探讨基于风险的MDP和贝尔曼方程对偶

报道来源 [3]

  1. arXiv cs.LG TIER_1 English(EN) · Kavya Ravichandran ·

    顺序决策和知识社会学中的算法方法

    arXiv:2607.20636v1 Announce Type: cross Abstract: As humans, we face many decisions that require us to choose between sticking to something and giving up. This thesis uses algorithmic tools to derive insights about such decision-making problems in theoretical models, studying bot…

  2. arXiv cs.AI TIER_1 English(EN) · Irmaan (Mohammad), Mirzanejad, Nadjet Bourdache, Abdel-Illah Mouaddib ·

    风险下的长期序贯决策制定

    arXiv:2607.19914v1 Announce Type: new Abstract: We study finite-horizon MDP planning under \emph{root-based} (resolute) risk objectives that apply a rank-dependent functional to the distribution of total returns. Such objectives are non-linear in the return distribution and gener…

  3. arXiv cs.AI TIER_1 English(EN) · Fernando E. Rosas, David Hyland, Daniel Polani ·

    序列决策中的广义贝尔曼递推和三个对偶

    arXiv:2607.18077v1 Announce Type: cross Abstract: What gives the Bellman equation its form? We show that the recursive properties of optimal value functions follow from three conditions: that the dynamics decomposes through sufficient statistics, that the return decomposes recurs…