Policy Evaluation
PulseAugur coverage of Policy Evaluation — every cluster mentioning Policy Evaluation across labs, papers, and developer communities, ranked by signal.
-
New algorithm tackles high variance in reinforcement learning policy evaluation
Researchers have developed a new double-loop gradient-based algorithm to address high variance in reinforcement learning policy evaluation. This method specifically tackles uncertainties in transition functions, which o…
-
DoTime generator enhances causal inference benchmarks for time series
Researchers have developed DoTime, a new synthetic benchmark generator designed to address the limitations of existing tools in evaluating causal inference for time series data. This open-source Python package aims to i…
-
New minimax PAC bounds for learning in exogenous contextual MDPs
Researchers have developed new minimax PAC bounds for learning in exogenous contextual Markov decision processes (MDPs). The study focuses on tabular discounted MDPs with exogenous, i.i.d. contexts that can influence re…
-
New research explores Bellman residual minimization for control tasks in reinforcement learning
This paper introduces foundational results for Bellman residual minimization applied to policy optimization in Markov decision problems. While dynamic programming is more common, Bellman residual minimization offers adv…