Researchers have developed a new estimator called Variance Optimal-CEM (VOCEM) for off-policy evaluation in contextual bandit policies. This method aims to reduce the high variance associated with action-level importance weighting by interpolating between existing estimators, OffCEM and Doubly Robust (DR). VOCEM selects an interpolation coefficient to minimize variance, resulting in an estimator that is no more variable than either OffCEM or DR. Experiments on synthetic data and benchmarks demonstrate that VOCEM outperforms both OffCEM and DR across various conditions, showing improved stability and empirical robustness. AI
IMPACT Improves stability and robustness in evaluating contextual bandit policies, potentially leading to more reliable reinforcement learning systems.
RANK_REASON The cluster contains a research paper detailing a new estimator for off-policy evaluation. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →