Contextual Bandits
PulseAugur coverage of Contextual Bandits — every cluster mentioning Contextual Bandits across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New algorithm enforces safety constraints in continuous action contextual bandits
Researchers have developed a new algorithm called High-Probability Constrained UCB for contextual bandit problems with continuous actions. This algorithm addresses safety concerns by enforcing high-probability constrain…
-
New BC-ICL method uses foundation models for efficient contextual bandits
Researchers have developed a new method called BC-ICL for contextual bandits, which leverages pre-trained tabular foundation models for more efficient personalization. This approach uses in-context learning with bootstr…
-
New research tackles cross-domain OPE/L for contextual bandits
A new research paper introduces Cross-Domain Off-Policy Evaluation and Learning (OPE/L) for contextual bandits, addressing limitations in existing methods. The proposed approach allows for the evaluation and learning of…
-
New Thompson Sampling methods tackle non-stationary and private contextual bandits
Two new research papers introduce novel approaches to Thompson sampling for contextual bandits. One paper, "Flow-Corrected Thompson Sampling for Non-Stationary Contextual Bandits," proposes a Bayesian method that reuses…
-
New framework optimizes risk-aware policy learning from logged data
Researchers have developed a new framework for risk-aware offline policy learning, which is essential for making decisions in high-stakes situations where real-time interaction is not possible. This approach allows for …