This paper investigates what aspects of a Markov decision process (MDP) can be identified solely from optimal actions, rather than direct observation of transition probabilities or Q-values. The research focuses on the identifiability of transition kernels under discounted MDPs, exploring how different forms of rewards (state-action, state-only, or state-action-next_state) influence what can be learned about the underlying dynamics. The findings indicate that knowing optimal actions for all state-action rewards is insufficient to uniquely determine transition probabilities, revealing an n(n-1)-dimensional family of kernels that yield the same optimal actions. AI
IMPACT This research contributes to a deeper theoretical understanding of reinforcement learning dynamics, potentially informing future algorithm development.
RANK_REASON Academic paper on a theoretical aspect of reinforcement learning. [lever_c_demoted from research: ic=1 ai=1.0]
- Deterministic policy gradient algorithms for semi‐Markov decision processes
- Letcher et al.
- Markov decision process
- Next-state functions for finite-state vector quantization
- policy
- Q-values
- Rewards and Fairies
- state-action rewards
- STATE REWARDS FOR MEDICAL DISCOVERIES
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →