Researchers have developed a new method for model-based reinforcement learning in stochastic environments that accounts for predictive uncertainty. This approach uses a quadratic action-value parameterization to simplify the Bellman backup, enabling analytic propagation of predictive mean and covariance. When applied with a Gaussian transition model and a radial-basis value function, the method results in a closed-form backup that reduces target variance and provides well-calibrated uncertainty estimates in continuous control tasks. AI
IMPACT This research offers a more principled framework for planning with learned distribution models in reinforcement learning, potentially improving performance in stochastic environments.
RANK_REASON The item is an academic paper detailing a new method for reinforcement learning. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- Bellman backup
- CatalyzeX Code Finder for Papers
- Connected Papers
- CORE Recommender
- DagsHub
- Gaussian function
- Gotit.pub
- Hugging Face
- IArxiv Recommender
- Influence Flower
- Litmaps
- ScienceCast
- scite Smart Citations
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →