OpenAI Gym
PulseAugur coverage of OpenAI Gym — every cluster mentioning OpenAI Gym across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
New STITCH-OPE framework uses guided diffusion for off-policy evaluation
Researchers have developed STITCH-OPE, a novel framework for off-policy evaluation (OPE) that utilizes guided diffusion models. This method is designed to handle high-dimensional, long-horizon problems in fields like ro…
-
New research explores robust, adaptive, and structured reinforcement learning techniques · 10 sources tracked
Multiple research papers published on arXiv explore advancements in reinforcement learning (RL) techniques. One paper unifies regularization-based methods for robust deep RL against adversarial perturbations, proposing …
-
SPADE framework uses LLM to design adaptive training environments for self-improvement
Researchers have developed SPADE, a novel self-play reinforcement learning framework where a single large language model acts as both an environment designer and a reasoning agent. The environment designer creates adapt…
-
New curriculum generation framework enhances autonomous agent navigation policies
Researchers have developed a new framework for generating curricula in structured parametric environments to enhance the robustness of navigation policies for autonomous agents. This method uses unidirectional gradient-…
-
New AI method uses LLMs to initialize reinforcement learning agents
Researchers have introduced ProDVI, a novel framework designed to enhance the sample efficiency of deep reinforcement learning agents. ProDVI utilizes large language models to generate Python code that hypothesizes envi…
-
AI boxing benchmark tests LLM decision speed and strategy
A developer has created an autonomous boxing benchmark to test AI decision-making, adaptability, and strategy. The benchmark simulates a fight with street rules, where AI models receive data about the match and can util…
-
LLMs explore preference alignment and failure mitigation techniques
Researchers are exploring new methods for aligning large language models (LLMs) with human preferences and mitigating specific failure modes. One approach uses Direct Preference Optimization (DPO) to reduce text degener…
-
Researchers fix synthetic data failures in reinforcement learning policy optimization
Researchers have identified and addressed algorithmic failures in Model-Based Policy Optimization (MBPO), a technique used in reinforcement learning. The study found that MBPO can underperform compared to other methods …
-
New interpretable experiential learning model shows promise for reinforcement learning
Researchers have introduced a novel interpretable experiential learning model that utilizes state history and global feedback to construct a behavioral model. This model represents learning as a transition graph between…