PulseAugur
EN
LIVE 04:18:52

New Kernel-WIS estimator improves off-policy evaluation for contextual bandits

Researchers have introduced Kernel-WIS, a new estimator for off-policy evaluation in contextual bandits. This method utilizes offline data and is designed to be asymptotically consistent. Kernel-WIS aims to outperform existing baselines by combining properties of vanilla weighted importance sampling and vanilla importance sampling, particularly in scenarios with misspecified behavior policies. AI

IMPACT This research could lead to more accurate and efficient methods for training reinforcement learning agents in contextual bandit settings.

RANK_REASON The cluster contains a research paper published on arXiv detailing a new algorithmic estimator.

Read on arXiv cs.LG →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

New Kernel-WIS estimator improves off-policy evaluation for contextual bandits

COVERAGE [2]

  1. arXiv cs.LG TIER_1 English(EN) · Joshua Spear, Matthieu Komorowski, Rebecca Pope, Neil J Sebire, Erica E. M. Moodie ·

    Kernel weighted importance sampling for off-policy evaluation in contextual bandits

    arXiv:2607.15067v1 Announce Type: new Abstract: This article presents a novel estimator for performing off-policy evaluation using only offline data for contextual bandits. The proposed estimator, Kernel-WIS is demonstrated to be asymptotically consistent and to empirically outpe…

  2. arXiv cs.LG TIER_1 English(EN) · Erica E. M. Moodie ·

    Kernel weighted importance sampling for off-policy evaluation in contextual bandits

    This article presents a novel estimator for performing off-policy evaluation using only offline data for contextual bandits. The proposed estimator, Kernel-WIS is demonstrated to be asymptotically consistent and to empirically outperform strong baselines (including vanilla weight…