Offline Reinforcement Learning
PulseAugur coverage of Offline Reinforcement Learning — every cluster mentioning Offline Reinforcement Learning across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
New research advances world models for embodied AI, focusing on robot behavior and deployment
Researchers are exploring advancements in embodied intelligence, focusing on "world models" that connect perception and decision-making for robots. Papers discuss frameworks for classifying these models from "plausible"…
-
New RL Policy Slashes Design Rule Violations by 92% in Chip Routing
Researchers have developed a history-aware offline reinforcement learning policy to address routing bottlenecks in physical design, particularly for complex, dense layouts. This new policy utilizes a lightweight LSTM ar…
-
Offline RL for Stroke Treatment Overstated Due to Confounding, Study Finds
A new research paper published on arXiv and highlighted by Hugging Face critically evaluates offline reinforcement learning (RL) methods used in stroke treatment. The study found that standard evaluation techniques, lik…
-
Bayesian Flow Networks unify discrete and continuous offline RL
Researchers have introduced Bayesian Flow Networks (BFNs) as a unified framework for offline reinforcement learning (RL) and trajectory planning. This new approach, termed BFN-RL, can natively handle both discrete and c…
-
New Delta Framework Detects Safety and Optimality Bugs in Deep RL Agents
Researchers have developed a new framework called Delta for identifying both safety and optimality bugs in Deep Reinforcement Learning (DRL) agents. This framework uses a two-phase approach: first, it evaluates the agen…
-
AI discovers high-quality chess puzzles using reinforcement learning
Researchers have developed a method using offline reinforcement learning to discover high-quality chess puzzles from a massive dataset of 1.5 billion puzzle-solving histories. This approach aims to improve the pedagogic…
-
VLMs used to annotate video game data for AI agent training · 2 sources tracked
Researchers are exploring the use of Vision Language Models (VLMs) to annotate video game data for training AI agents. The first paper investigates VLMs' ability to generate reward signals from game sequences, identifyi…
-
New method CSDG enhances offline reinforcement learning
Researchers have introduced Convex Hull Neighborhood Smooth Dual Generalization (CSDG), a novel method for offline reinforcement learning. CSDG addresses the issue of amplified estimation errors in out-of-distribution a…
-
Research paper questions robustness of offline RL for treatment recommendations
A new research paper published on arXiv investigates the effectiveness of covariate balance diagnostics in long time horizon Markov decision processes, particularly within the context of offline reinforcement learning f…
-
New FORE method improves offline reinforcement learning evaluation
Researchers have introduced Fitted Occupancy-Ratio Evaluation (FORE), a novel method for estimating occupancy ratios in offline reinforcement learning. This technique characterizes the discounted occupancy ratio through…
-
AI research explores structural generalization in RL and NLP · 2 sources tracked
Two new research papers explore different facets of generalization in AI models. The first paper, focusing on offline reinforcement learning, argues that the structure of pessimism in datasets is more critical for gener…
-
New dataset Insulin4RL enables offline reinforcement learning with irregular clinical data
Researchers have introduced Insulin4RL, a new dataset designed for offline reinforcement learning in healthcare settings. This dataset, derived from MIMIC-IV, contains over 375,000 decisions from 12,209 intensive care u…
-
New framework refines offline RL trajectories using counterfactual flows
Researchers have introduced a new framework called counterfactual transport flows for offline reinforcement learning. This method aims to improve decision-making policies using only logged historical data, without extra…
-
New benchmark standardizes offline RL for nuclear fusion plasma control
Researchers have introduced RL4F, a new benchmark designed to standardize the evaluation of offline reinforcement learning for plasma control in nuclear fusion. This benchmark utilizes historical data from the DIII-D to…
-
New TrojanTO attack targets trajectory optimization models in RL
Researchers have developed TrojanTO, a novel method for executing action-level backdoor attacks against trajectory optimization (TO) models used in offline reinforcement learning. Unlike previous reward-manipulation att…
-
New bootstrap method enhances offline reinforcement learning analysis
Researchers have developed a new model-based bootstrap method for controlled Markov chains, particularly useful in offline reinforcement learning scenarios where the data-generating policy is unknown. This technique est…
-
New ME-AM framework enhances offline RL with entropy maximization
Researchers have introduced Maximum Entropy Adjoint Matching (ME-AM), a new framework designed to improve offline reinforcement learning. This method addresses limitations in existing approaches, such as popularity bias…
-
New Q-Ising method optimizes dynamic treatment allocation on networks
Researchers have developed Q-Ising, a novel three-stage pipeline for dynamic treatment allocation in networks. This method integrates network structure with dynamic treatment strategies, addressing limitations of existi…
-
New AdamO optimizer enhances stability and performance in offline RL
Researchers have introduced AdamO, a novel optimizer designed to enhance stability in offline reinforcement learning. This new optimizer addresses the issue of 'collapse,' where errors in temporal-difference updates can…