PulseAugur
EN
LIVE 06:03:25

Research paper details convergence of optimistic policy iteration for stochastic shortest path problem

This paper, submitted to arXiv's Computer Science > Machine Learning section, presents convergence results for a specific optimistic policy iteration algorithm applied to the stochastic shortest path problem. The research, authored by Yuanlong Chen, details the use of Monte Carlo and TD(lambda) methods for policy evaluation, assuming the termination state is eventually reached with certainty. AI

RANK_REASON The cluster contains an academic paper on a machine learning topic. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.LG →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Research paper details convergence of optimistic policy iteration for stochastic shortest path problem

COVERAGE [1]

  1. arXiv cs.LG TIER_1 English(EN) · Yuanlong Chen ·

    On the convergence of optimistic policy iteration for stochastic shortest path problem

    arXiv:1808.08763v3 Announce Type: replace Abstract: In this paper, we prove some convergence results of a special case of optimistic policy iteration algorithm for stochastic shortest path problem. We consider both Monte Carlo and $TD(\lambda)$ methods for the policy evaluation s…