PulseAugur
实时 01:27:22
English(EN) Kernel weighted importance sampling for off-policy evaluation in contextual bandits

新的Kernel-WIS估计器改进了上下文老虎机中的离策略评估

研究人员推出了一种用于上下文老虎机中离策略评估的新估计器Kernel-WIS。该方法利用离线数据,并且设计为渐近一致的。Kernel-WIS旨在通过结合标准加权重要性采样和标准重要性采样的特性,尤其是在行为策略错误指定的情况下,优于现有基线。 AI

影响 这项研究可能为上下文老虎机设置中强化学习代理的训练带来更准确、更有效的方法。

排序理由 该集群包含一篇在arXiv上发表的详细介绍新算法估计器的研究论文。

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

新的Kernel-WIS估计器改进了上下文老虎机中的离策略评估

报道来源 [2]

  1. arXiv cs.LG TIER_1 English(EN) · Joshua Spear, Matthieu Komorowski, Rebecca Pope, Neil J Sebire, Erica E. M. Moodie ·

    上下文赌徒中离策略评估的核加权重要性采样

    arXiv:2607.15067v1 Announce Type: new Abstract: This article presents a novel estimator for performing off-policy evaluation using only offline data for contextual bandits. The proposed estimator, Kernel-WIS is demonstrated to be asymptotically consistent and to empirically outpe…

  2. arXiv cs.LG TIER_1 English(EN) · Erica E. M. Moodie ·

    上下文赌徒中离策略评估的核加权重要性采样

    This article presents a novel estimator for performing off-policy evaluation using only offline data for contextual bandits. The proposed estimator, Kernel-WIS is demonstrated to be asymptotically consistent and to empirically outperform strong baselines (including vanilla weight…