A new research paper introduces Cross-Domain Off-Policy Evaluation and Learning (OPE/L) for contextual bandits, addressing limitations in existing methods. The proposed approach allows for the evaluation and learning of new policies using historical data from both the target domain and other related domains. This novel formulation aims to enhance OPE/L performance in challenging scenarios such as limited data, deterministic logging policies, and the introduction of new actions, which are common in fields like personalized medicine and content recommendations. AI
IMPACT This research could improve the effectiveness of personalized systems by enabling better policy evaluation and learning in domains with limited or complex historical data.
RANK_REASON The item is an academic paper submitted to arXiv. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX Code Finder for Papers
- Contextual Bandits
- CORE Recommender
- Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits
- Cross-Domain OPE/L
- DagsHub
- Gotit.pub
- Hugging Face
- IArxiv Recommender
- Influence Flower
- Off-Policy Evaluation and Learning for External Validity under a Covariate Shift
- OPE/L
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →