PulseAugur
EN
LIVE 08:17:23

Research explores identifiability of transition kernels in discounted MDPs

This paper investigates what aspects of a Markov decision process (MDP) can be identified solely from optimal actions, rather than direct observation of transition probabilities or Q-values. The research focuses on the identifiability of transition kernels under discounted MDPs, exploring how different forms of rewards (state-action, state-only, or state-action-next_state) influence what can be learned about the underlying dynamics. The findings indicate that knowing optimal actions for all state-action rewards is insufficient to uniquely determine transition probabilities, revealing an n(n-1)-dimensional family of kernels that yield the same optimal actions. AI

IMPACT This research contributes to a deeper theoretical understanding of reinforcement learning dynamics, potentially informing future algorithm development.

RANK_REASON Academic paper on a theoretical aspect of reinforcement learning. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.LG →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Research explores identifiability of transition kernels in discounted MDPs

COVERAGE [1]

  1. arXiv cs.LG TIER_1 English(EN) · Neal Batra ·

    From Optimal Actions to World Models: Identifiability of Transition Kernels in Discounted MDPs

    arXiv:2608.07301v1 Announce Type: new Abstract: We study what can be recovered about the transition probabilities of a Markov decision process from optimal actions alone. This is closely related to the inverse problem considered by Letcher et al., who ask when the dynamics can be…