一篇新论文提出了玻尔兹曼理性(一种常见的随机选择模型)的公理化表征,该模型在强化学习中得到广泛应用。该研究区分了选择中的随机性与环境中的偶然性,并提出通过将独立性公理限制在环境彩票上,可以唯一地推导出玻尔兹曼策略及其相关的软贝尔曼方程。该框架为智能体设计提供了规范性评估,并阐明了无关选项独立性公理适用的条件。 AI
影响 为理解强化学习中智能体的决策过程提供了理论框架。
排序理由 该集群包含一篇发表在arXiv上的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
- Boltzmann policy
- Boltzmann Rationality
- economics
- hard Bellman equation
- Independence
- information theory
- Markov decision process
- reinforcement learning
- soft Bellman equation
- softmax policy
- von Neumann–Morgenstern utility theorem
AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →