Markov decision process
PulseAugur coverage of Markov decision process — every cluster mentioning Markov decision process across labs, papers, and developer communities, ranked by signal.
- instance of CatalyzeX 90%
- instance of deep reinforcement learning 90%
- used by Deep Q-Network 90%
- instance of ScienceCast 90%
- used by ScienceCast 70%
- used by Gotit.pub 70%
- instance of alphaXiv 70%
- instance of Influence Flower 70%
- used by deep reinforcement learning 70%
- instance of Gotit.pub 70%
- used by dynamic programming 70%
- affiliated with Deep Q-Network 70%
18 day(s) with sentiment data
-
New framework analyzes optimal policies in Restart POMDPs
Researchers have developed a new framework for analyzing Restart POMDPs (Partially Observable Markov Decision Processes) on general Borel state spaces. This framework reduces the problem to a fully observed MDP by using…
-
PressureMesh system estimates 3D human poses using multi-device pressure data
Researchers have developed PressureMesh, a novel system for estimating 3D human poses using pressure data from multiple devices. The system utilizes an end-to-end network called MDP-Net, which incorporates a Mixture of …
-
New path-dependent inference method enhances discrete object sampling
Researchers have developed a new method for sampling compositional and discrete objects from unnormalized posterior distributions. This approach, termed path-dependent discrete amortized inference, enhances existing Mar…
-
Soft Actor-Critic enhances heat pump control, reducing wear and improving efficiency
Researchers have developed a new reinforcement learning approach using Soft Actor-Critic (SAC) to improve the control of inverter-driven heat pumps. This method aims to reduce compressor wear by minimizing on-off cyclin…
-
New algorithm DTOA tackles team games with inaccurate supervisor beliefs
Researchers have developed a new algorithm called the Distributed Team Orchestrating Algorithm (DTOA) to address challenges in zero-sum potential team games where agents rely on potentially inaccurate belief information…
-
Research explores identifiability of transition kernels in discounted MDPs
This paper investigates what aspects of a Markov decision process (MDP) can be identified solely from optimal actions, rather than direct observation of transition probabilities or Q-values. The research focuses on the …
-
New framework tackles hidden Byzantine attacks in multi-agent systems
Researchers have developed a new theoretical framework and algorithm for online security learning in cooperative multi-agent systems facing hidden Byzantine attacks. The study identifies that an attacker's information a…
-
New frameworks advance Markov Decision Processes for reinforcement learning · 2 sources tracked
Two new arXiv papers introduce advanced frameworks for Markov Decision Processes (MDPs), a key tool in reinforcement learning. The first paper, GRASP-MDP, addresses challenges in offline reinforcement learning by separa…
-
New AI optimizer enhances military asset placement against adversarial threats
A new research paper introduces an advanced optimization engine for military asset placement, addressing the critical and previously unsolved problem of pre-commitment posture. The proposed Composite Expected Value (CEV…
-
Multi-agent RL enhances UAV deployment and communication in sparse networks
Two new research papers explore the application of multi-agent reinforcement learning for optimizing the deployment and communication of unmanned aerial vehicles (UAVs). The first paper introduces a framework for decent…
-
New POMDP framework separates intent from action in noisy social dilemmas
Researchers have developed a new framework for understanding intentions in social dilemmas where actions are subject to noise. By using a Partially Observable MDP (POMDP) formulation, the model can distinguish between a…
-
New MDP Algorithm Integrates Q-Value Predictions for Enhanced Robustness
Researchers have developed a new framework for Markov Decision Processes (MDPs) that improves upon traditional methods by incorporating Q-value predictions. This approach moves beyond treating machine-learned advice as …
-
New causal abstraction technique improves MDP scalability
Researchers have developed a new property-driven causal abstraction technique for Markov Decision Processes (MDPs) to address scalability challenges. This method leverages causal relations over state variable predicates…
-
SymmGrid framework accelerates on-robot learning with parallelized symmetries
Researchers have developed SymmGrid, a new framework designed to significantly accelerate on-robot learning for deep reinforcement policies. By leveraging parallelized symmetries within a Markov Decision Process, SymmGr…
-
New research offers finite-time convergence guarantees for Natural Policy Gradient algorithms
A new research paper published on arXiv provides the first finite-time convergence guarantees for Natural Policy Gradient (NPG) algorithms in finite-horizon Markov Decision Processes. The study analyzes NPG under both c…
-
New VRDQ algorithm enables faster decentralized reinforcement learning
Researchers have developed a new decentralized reinforcement learning algorithm called VRDQ. This algorithm is designed for scenarios where multiple agents interact with the same Markov Decision Process and can share in…
-
DynaMark framework uses RL for dynamic watermarking in industrial MTCs
Researchers have developed DynaMark, a novel reinforcement learning framework designed to enhance security in industrial Machine Tool Controllers (MTCs) within Industry 4.0 environments. This system addresses vulnerabil…
-
Active Inference framed as convex MDP, unifying with reinforcement learning
A new paper frames Active Inference (AIF) as a convex Markov Decision Process (MDP), suggesting a way to unify it with modern reinforcement learning (RL) techniques. The research posits that minimizing expected free ene…
-
New simulator Eutopia tackles long-term AI fairness in credit lending
Researchers have developed Eutopia, a simulator designed to evaluate long-term fairness in AI-driven decision-making, specifically within a credit lending context. This simulator addresses limitations in existing fairne…
-
New drone trajectory planner detects and counters ID spoofing attacks
Researchers have developed a new trajectory planning framework for small unmanned aerial systems (UAS) that accounts for potential Remote Identification (RID) spoofing attacks. Unlike existing methods that trust RID bro…