PulseAugur
EN
LIVE 19:24:04
ENTITY Offline Reinforcement Learning

Offline Reinforcement Learning

PulseAugur coverage of Offline Reinforcement Learning — every cluster mentioning Offline Reinforcement Learning across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
3
13 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
3
13 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

3 day(s) with sentiment data

RECENT · PAGE 1/1 · 13 TOTAL
  1. RESEARCH · CL_187235 ·

    VLMs used to annotate video game data for AI agent training · 2 sources tracked

    Researchers are exploring the use of Vision Language Models (VLMs) to annotate video game data for training AI agents. The first paper investigates VLMs' ability to generate reward signals from game sequences, identifyi…

  2. TOOL · CL_183289 ·

    New method CSDG enhances offline reinforcement learning

    Researchers have introduced Convex Hull Neighborhood Smooth Dual Generalization (CSDG), a novel method for offline reinforcement learning. CSDG addresses the issue of amplified estimation errors in out-of-distribution a…

  3. RESEARCH · CL_147426 ·

    Research paper questions robustness of offline RL for treatment recommendations

    A new research paper published on arXiv investigates the effectiveness of covariate balance diagnostics in long time horizon Markov decision processes, particularly within the context of offline reinforcement learning f…

  4. RESEARCH · CL_128341 ·

    New FORE method improves offline reinforcement learning evaluation

    Researchers have introduced Fitted Occupancy-Ratio Evaluation (FORE), a novel method for estimating occupancy ratios in offline reinforcement learning. This technique characterizes the discounted occupancy ratio through…

  5. RESEARCH · CL_123086 ·

    AI research explores structural generalization in RL and NLP · 2 sources tracked

    Two new research papers explore different facets of generalization in AI models. The first paper, focusing on offline reinforcement learning, argues that the structure of pessimism in datasets is more critical for gener…

  6. TOOL · CL_100180 ·

    New dataset Insulin4RL enables offline reinforcement learning with irregular clinical data

    Researchers have introduced Insulin4RL, a new dataset designed for offline reinforcement learning in healthcare settings. This dataset, derived from MIMIC-IV, contains over 375,000 decisions from 12,209 intensive care u…

  7. TOOL · CL_80057 ·

    New framework refines offline RL trajectories using counterfactual flows

    Researchers have introduced a new framework called counterfactual transport flows for offline reinforcement learning. This method aims to improve decision-making policies using only logged historical data, without extra…

  8. TOOL · CL_79775 ·

    New benchmark standardizes offline RL for nuclear fusion plasma control

    Researchers have introduced RL4F, a new benchmark designed to standardize the evaluation of offline reinforcement learning for plasma control in nuclear fusion. This benchmark utilizes historical data from the DIII-D to…

  9. TOOL · CL_58992 ·

    New TrojanTO attack targets trajectory optimization models in RL

    Researchers have developed TrojanTO, a novel method for executing action-level backdoor attacks against trajectory optimization (TO) models used in offline reinforcement learning. Unlike previous reward-manipulation att…

  10. RESEARCH · CL_29303 ·

    New bootstrap method enhances offline reinforcement learning analysis

    Researchers have developed a new model-based bootstrap method for controlled Markov chains, particularly useful in offline reinforcement learning scenarios where the data-generating policy is unknown. This technique est…

  11. TOOL · CL_21970 ·

    New ME-AM framework enhances offline RL with entropy maximization

    Researchers have introduced Maximum Entropy Adjoint Matching (ME-AM), a new framework designed to improve offline reinforcement learning. This method addresses limitations in existing approaches, such as popularity bias…

  12. RESEARCH · CL_21748 ·

    New Q-Ising method optimizes dynamic treatment allocation on networks

    Researchers have developed Q-Ising, a novel three-stage pipeline for dynamic treatment allocation in networks. This method integrates network structure with dynamic treatment strategies, addressing limitations of existi…

  13. TOOL · CL_16081 ·

    New AdamO optimizer enhances stability and performance in offline RL

    Researchers have introduced AdamO, a novel optimizer designed to enhance stability in offline reinforcement learning. This new optimizer addresses the issue of 'collapse,' where errors in temporal-difference updates can…