PulseAugur
实时 07:49:10
English(EN) Unified continuous-time q-learning for mean-field game and mean-field control problems

统一的q学习框架用于均值场博弈和控制问题

研究人员开发了一个统一的连续时间q学习框架,用于均值场博弈和控制问题。这种方法被称为解耦的Iq函数,建立了一个鞅刻画,它作为均值场博弈(MFG)和均值场控制(MFC)场景的通用策略评估规则。所提出的算法即使在环境模拟器无法直接访问种群分布的情况下也有效,它基于代表性代理的状态值更新种群分布。该框架的效用通过在LQ框架内外的应用得到证明,展示了其在MFG和MFC学习任务中的效率。 AI

影响 引入了一种新颖的统一q学习方法,用于复杂的博弈和控制问题,可能推动强化学习的应用。

排序理由 该集群包含一篇在arXiv上发表的研究论文,详细介绍了新的理论框架和算法。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

统一的q学习框架用于均值场博弈和控制问题

报道来源 [1]

  1. arXiv cs.LG TIER_1 English(EN) · Xiaoli Wei, Xiang Yu, Fengyi Yuan ·

    统一连续时间q学习用于均场博弈和均场控制问题

    arXiv:2407.04521v3 Announce Type: replace-cross Abstract: This paper studies the continuous-time q-learning in mean-field jump-diffusion models in a setting where the environment simulator does not provide direct access to the population distribution. We propose the integrated q-…