Markov decision processes: a tool for sequential decision making under uncertainty
PulseAugur coverage of Markov decision processes: a tool for sequential decision making under uncertainty — every cluster mentioning Markov decision processes: a tool for sequential decision making under uncertainty across labs, papers, and developer communities, ranked by signal.
10 day(s) with sentiment data
-
New geometric theory analyzes decision boundaries in structured MDPs
This paper introduces a novel geometric theory for analyzing optimal policies in structured Markov Decision Processes (MDPs). It proposes that the geometry of the decision boundary, rather than the size of the state spa…
-
New framework for social laws in stochastic multi-agent systems unveiled
This paper introduces a new framework for social laws in multi-agent systems operating in stochastic environments. It extends previous work from deterministic settings to reward-based scenarios, proposing a method for d…
-
New LP-based algorithm offers stronger policies for Submodular MDPs
Researchers have developed a new algorithm for solving Submodular Markov Decision Processes (MDPs), a type of sequential decision-making problem with generalized reward functions. The algorithm, based on Linear Programm…
-
New Windowed A-K-MDP algorithm improves conservation decision-making
Researchers have introduced Windowed A-K-MDP, an enhanced algorithm for Markov decision processes (MDPs) designed to improve decision-making in areas like biodiversity conservation. This new method addresses limitations…
-
New Bellman Equation Targets Plasticity in Reinforcement Learning
Researchers have introduced a new Bellman optimality equation specifically designed to optimize plasticity in continual reinforcement learning. This work builds upon a prior formalization that reframed the stability-pla…
-
New algorithm enables safe learning in irreversible environments
Researchers have developed a novel learning algorithm designed for agents operating in environments with irreversible dynamics, where mistakes cannot be undone. This algorithm allows agents to request assistance from a …
-
Open-source ZGCM-1 model achieves high efficiency in math and agentic search
Researchers have introduced ZGCM-1, a 7B parameter foundation model designed for mathematical reasoning and agentic search. The model leverages an efficient training recipe that combines architectural innovations like i…
-
New algorithm \Algname enhances Monte Carlo Tree Search for stochastic environments
Researchers have developed a new Monte Carlo Tree Search (MCTS) algorithm called \Algname, specifically designed for continuous and stochastic Markov Decision Processes (MDPs). This novel approach integrates a power mea…
-
New risk-averse Q-learning method for robot navigation
Researchers have developed a novel approach to risk-averse reinforcement learning for complex decision-making tasks. This method, termed Mini-Batch Risk-Averse Deep Q-Learning, addresses the challenge of estimating tran…
-
New RCSD method optimizes multi-agent coordination under communication limits
Researchers have developed a new method called Reachability-Certified Subteam Decomposition (RCSD) for multi-agent systems operating under communication constraints. This technique aims to optimize coordination by consi…
-
New federated RL algorithm minimizes communication costs
Researchers have introduced Fed-LSVI, a novel federated algorithm designed for online reinforcement learning with linear function approximation. This algorithm addresses the communication and privacy challenges inherent…
-
New benchmarks tackle LLM agent complexity in healthcare
Two new research papers propose advanced benchmarking protocols for large language model (LLM) agents in healthcare settings. The first paper introduces an episode-level evaluation protocol that separates evidence acros…
-
New method uses model checking to test LLM explanations
Researchers have developed a novel method for automatically testing the accuracy of explanations generated by large language models (LLMs) when used to interpret sequential decision-making policies. This approach utiliz…
-
GFlowNets applied to solve complex combinatorial optimization problems
Researchers have developed a novel approach using GFlowNets to tackle complex combinatorial optimization problems, which are often too difficult for traditional algorithms. The method involves designing specific Markov …
-
New categorizer automata improve AI data binning efficiency
Researchers have introduced a new type of automaton called the categorizer automaton, designed for AI systems to categorize continuous data into discrete bins. This new automaton generalizes comparator automata and offe…
-
New framework offers quantitative analysis for robust Markov Decision Processes
This paper introduces a quantitative analysis framework for robust Markov Decision Processes (RMDPs) with $ω$-regular objectives. The research extends previous qualitative analyses by solving for the exact quantitative …
-
New method tackles complexity in multi-agent decision-making
Researchers have introduced a novel approach to address the exponential complexity of Decentralised Partially Observable Markov Decision Processes (DecPOMDPs) in multi-agent systems. The paper proposes shifting focus fr…
-
New bounds established for constrained average-reward MDPs
Researchers have established near-optimal sample complexity bounds for constrained average-reward Markov decision processes (CAMDPs) under a generative model. The proposed model-based algorithm achieves sample complexit…
-
New algorithms tackle decentralized multi-player reinforcement learning
Researchers have developed new algorithms for decentralized multi-player reinforcement learning in episodic Markov Decision Processes (MDPs) with information asymmetry. The proposed methods, mQ-learning, mQ-learning-int…
-
New sub-quadratic method improves bisimulation metric computation for MDPs
Researchers have developed a novel sub-quadratic method for calculating bisimulation metrics in Markov decision processes (MDPs). This new approach utilizes approximate nearest neighbor (ANN) indexing to efficiently sel…