Researchers have developed new online learning algorithms for policy evaluation in Markov decision processes (MDPs) that incorporate dynamic utility-based shortfall risk (UBSR) measures. The proposed UBSR-TD algorithm and its variants are designed for efficient use with linear function approximation, establishing conditions for almost sure convergence. These methods adapt existing risk-neutral policy evaluation algorithms by modifying the temporal-difference error with a specific loss function, and their practical utility is demonstrated through experiments, including an application to perishable inventory management under shelf-life uncertainty. AI
IMPACT Enhances risk-aware reinforcement learning capabilities for sequential decision-making problems.
RANK_REASON Academic paper detailing new algorithms for policy evaluation in MDPs. [lever_c_demoted from research: ic=1 ai=1.0]
- Markov decision processes
- Perishable Inventory Management Using GA-ANN and ICA-ANN
- reinforcement learning
- UbSRD: The Ubiquitin Structural Relational Database
- UBSR-TD
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →