Researchers have introduced Kernel-WIS, a new estimator for off-policy evaluation in contextual bandits. This method utilizes offline data and is designed to be asymptotically consistent. Kernel-WIS aims to outperform existing baselines by combining properties of vanilla weighted importance sampling and vanilla importance sampling, particularly in scenarios with misspecified behavior policies. AI
IMPACT This research could lead to more accurate and efficient methods for training reinforcement learning agents in contextual bandit settings.
RANK_REASON The cluster contains a research paper published on arXiv detailing a new algorithmic estimator.
- alphaXiv
- arXiv
- CatalyzeX Code Finder for Papers
- CORE Recommender
- DagsHub
- Gotit.pub
- Hugging Face
- IArxiv Recommender
- importance sampling
- Influence Flower
- Kernel-WIS
- ScienceCast
- Weighted importance sampling for off-policy learning with linear function approximation
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →