Markov decision process
PulseAugur coverage of Markov decision process — every cluster mentioning Markov decision process across labs, papers, and developer communities, ranked by signal.
- instance of alphaXiv 90%
- instance of CatalyzeX 90%
- instance of deep reinforcement learning 90%
- used by Deep Q-Network 90%
- used by alphaXiv 70%
- instance of ScienceCast 70%
- instance of Influence Flower 70%
- instance of CatalyzeX Code Finder for Papers 70%
- used by CatalyzeX 70%
- used by Gotit.pub 70%
- used by ScienceCast 70%
- instance of CORE Recommender 70%
7 day(s) with sentiment data
-
New AI control framework allows delegation to misaligned agents
Researchers have developed a new framework for controlling AI agents, particularly those that operate over long durations. The proposed method, termed 'k-robust coalitional alignment,' allows for delegation of authoriza…
-
AI framework cuts 5G energy use while preserving service levels
Researchers have developed a new AI-driven framework for energy saving in 5G networks that ensures service-level agreements (SLAs) are maintained. The system uses a stability-aware constrained reinforcement learning app…
-
New algorithm offers efficient solutions for Markov Decision Processes
Researchers have developed a novel algorithm for efficiently solving Markov Decision Process (MDP) problems that utilize function approximations. This new method is based on a linear programming reformulation and iterat…
-
New benchmark Sim2Signal tackles Sim-to-Real gap in traffic signal control
Researchers have introduced Sim2Signal, a new benchmark designed to systematically measure and evaluate methods for bridging the "Sim-to-Real gap" in traffic signal control using reinforcement learning. This gap, where …
-
New DTMDP model tackles lot-sizing with stochastic demand timing
Researchers have developed a discrete-time Markov decision process (DTMDP) model to address a multi-item capacitated lot-sizing problem with stochastic demand timing. This model accounts for factors like capacity compet…
-
New A-MADiff algorithm optimizes AI-generated content delivery in mobile networks
Researchers have developed A-MADiff, a novel multi-agent reinforcement learning algorithm designed to optimize task orchestration in mobile networks that host Artificial Intelligence-Generated Content (AIGC) services. T…
-
LLM AI agents bypass complex probability calculations despite theoretical frameworks
Large Language Models (LLMs) used in AI agents do not inherently calculate probabilities, despite theoretical frameworks like Partially Observable Markov Decision Processes (POMDPs) suggesting they should. While POMDPs …
-
New hierarchical RL framework enhances conversational agents
Researchers have developed a novel two-level hierarchical reinforcement learning (RL) framework called ToSCA for conversational agents. This approach bridges the gap between existing token-level or utterance-level RL me…
-
iScheduler uses RL to optimize large-scale resource allocation
Researchers have developed iScheduler, a new framework that uses reinforcement learning to optimize resource allocation for large-scale computing tasks. This approach models the Resource Investment Problem (RIP) as a Ma…
-
New analysis explores faster convergence for Natural Actor-Critic algorithms
Researchers have analyzed a single-loop, entropy-regularized Natural Actor-Critic algorithm, focusing on its convergence rates for unregularized objectives. The study explores two optimization regimes: Stochastic, using…
-
New framework enables parallel planning for multi-agent path finding
Researchers have developed a theoretical framework for parallel lifelong multi-agent path finding (L-MAPF) using group decentralized planning. The new Group Decentralized RHCR (GD-RHCR) framework builds upon the Rolling…
-
New research frames MDP planning as Bayesian inference over policies
Researchers have proposed a novel approach to Markov decision process (MDP) planning by framing it as a problem of Bayesian inference over policies. This conceptual shift treats the policy itself as a latent variable, w…
-
Pointer Networks with Q-Learning for Combinatorial Optimization
A research paper introduces the Pointer Q-Network (PQN), a novel neural architecture designed to improve sequence generation for combinatorial optimization tasks. The PQN integrates model-free Q-value approximation with…
-
RLCascadeRouter optimizes LLM routing with reinforcement learning
Researchers have developed RLCascadeRouter, a novel framework that optimizes query routing for large language models (LLMs) by treating it as a Markov decision process. This approach directly optimizes the performance-c…
-
AI advice channels can disempower humans by fostering reliance, study finds
A new research paper explores how AI advice channels can subtly disempower humans by fostering reliance. The study models the fraction of behavior that follows AI advice as a state within a Markov decision process, demo…
-
Offline RL optimizes sepsis treatment using MIMIC-IV data
Researchers have developed a novel approach using offline reinforcement learning to optimize the management of sepsis in intensive care units. By analyzing historical patient data from the MIMIC-IV database, the study m…
-
New AI framework combines reinforcement learning and symbolic heuristics for temporal planning
Researchers have developed a new framework that enhances temporal planning by integrating reinforcement learning with symbolic heuristics. This approach aims to improve the performance of AI planners by learning heurist…
-
New AI model generates state-of-the-art fashion outfits
Researchers have developed a new framework for generating fashion outfits, addressing the complexity of aesthetic compatibility and large search spaces. The proposed Unified Sequential Composition Model (USCM) formalize…
-
New framework analyzes optimal policies in Restart POMDPs
Researchers have developed a new framework for analyzing Restart POMDPs (Partially Observable Markov Decision Processes) on general Borel state spaces. This framework reduces the problem to a fully observed MDP by using…
-
PressureMesh system estimates 3D human poses using multi-device pressure data
Researchers have developed PressureMesh, a novel system for estimating 3D human poses using pressure data from multiple devices. The system utilizes an end-to-end network called MDP-Net, which incorporates a Mixture of …