PulseAugur
实时 07:46:41
English(EN) Online Inference in Distributional Temporal-Difference Learning

揭示了时序差分学习的新的统计推断方法

研究人员开发了用于时序差分学习中回报分布的在线统计推断的新方法。该研究证明了 Polak-Ruppert 平均估计量的根号T误差在 Cramér 空间中弱收敛于一个中心高斯随机元素。这项工作验证了平滑统计泛函(如方差和分位数)的 bootstrap 推断,并为非平滑泛函引入了局部渐近理论。 AI

影响 引入了适用于时序差分学习的新型统计推断技术,有可能提高模型的准确性和可靠性。

排序理由 该集群包含一篇在 arXiv 上发表的研究论文,详细介绍了时序差分学习的新统计方法。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv stat.ML 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

揭示了时序差分学习的新的统计推断方法

报道来源 [1]

  1. arXiv stat.ML TIER_1 English(EN) · Yang Peng, Liangyu Zhang ·

    在线分布时序差分学习中的在线推理

    arXiv:2608.14408v1 Announce Type: new Abstract: We study online statistical inference for functionals of the return distribution under a fixed policy. The return distribution is estimated by nonparametric distributional temporal-difference learning from a single Markov trajectory…