Proximal Policy Optimization
PulseAugur coverage of Proximal Policy Optimization — every cluster mentioning Proximal Policy Optimization across labs, papers, and developer communities, ranked by signal.
- instance of reinforcement learning 90%
- instance of deep reinforcement learning 90%
- instance of Pfadfinder und Pfadfinderinnen Österreichs 90%
- developed Advantage Actor-Critic 90%
- used by long short-term memory 90%
- used by MuJoCo 90%
- competes with Grpo 70%
- used by Grpo 70%
- used by large-language models 70%
- used by reinforcement learning from human feedback 70%
- developed Grpo 70%
- used by TD3 70%
- 2026-05-26 research_milestone A new method is proposed to stabilize reinforcement learning training by strategically dropping transitions. source
24 day(s) with sentiment data
-
New Evaluation-Conditioned Training method improves LLM generalization
Researchers have introduced Evaluation-Conditioned Training (ECT), a novel post-training framework designed to enhance the generalization capabilities of large language models (LLMs). This method aims to address limitat…
-
New DClamp-PPO algorithm enhances reinforcement learning by penalizing 'wrong' direction updates
Researchers have introduced Directional-Clamp PPO (DClamp-PPO), a novel algorithm designed to enhance the performance of Proximal Policy Optimization (PPO) in deep reinforcement learning. DClamp-PPO addresses a key limi…
-
AI agent achieves 97.5% collision avoidance in space simulations
Researchers have developed a reinforcement learning policy using Proximal Policy Optimization (PPO) to autonomously avoid collisions in space. This new approach aims to address the growing problem of orbital congestion …
-
Paper introduces adversarial training for robust RL policies
This paper, titled "Adversarial Latent-State Training for Robust Policies in Partially Observable Domains," introduces a new framework for reinforcement learning in partially observable environments. The authors propose…
-
New AI benchmark environment created for Dark Souls boss fights · 2 sources tracked
Researchers have developed the Dark Souls Learning Environment (DSLE), a platform designed to benchmark AI agents against the boss encounters in Dark Souls: Remastered. The environment presents 22 boss fights as challen…
-
New Dyna-style RL approach cuts quadrupedal robot training time
Researchers have developed a new Dyna-style framework that significantly improves the data efficiency of reinforcement learning (RL) for quadrupedal locomotion. By integrating a learned transition model to generate synt…
-
AI tackles hurricane disruption in freight routing with multi-agent RL
Researchers have developed a per-shipment multi-agent reinforcement learning approach for intermodal freight routing, specifically addressing disruptions from events like hurricanes. Their Independent PPO (IPPO) method,…
-
New ABC-GRPO algorithm enhances LLM training stability and performance
Researchers have introduced All-Quadrant Bounded Clipping GRPO (ABC-GRPO), a novel algorithm designed to improve the stability and generalizability of reinforcement learning for large language models. ABC-GRPO addresses…
-
Reinforcement learning method enhances cleaning robot path planning
A new research paper proposes an improved path planning method for cleaning robots using reinforcement learning. The method combines the Proximal Policy Optimization (PPO) algorithm with transfer learning, a 'detection …
-
AI agents learn 'guilt' from human neural data for prosocial behavior
Researchers have developed a method to calibrate artificial guilt signals for prosocial multi-agent reinforcement learning by analyzing human neural and behavioral data. Using fMRI data from 40 participants, they derive…
-
Multi-agent RL enhances UAV deployment and communication in sparse networks
Two new research papers explore the application of multi-agent reinforcement learning for optimizing the deployment and communication of unmanned aerial vehicles (UAVs). The first paper introduces a framework for decent…
-
Machine learning agent learns reactive play in Atari Breakout
A machine learning experiment successfully trained an agent to play Atari Breakout reactively, rather than relying on scripted actions. After 123 failed attempts using various reinforcement learning techniques, the brea…
-
SynAgent framework enables scalable cooperative humanoid manipulation
Researchers have introduced SynAgent, a novel framework designed to enhance cooperative humanoid manipulation capabilities. This system addresses data scarcity and coordination complexities by transferring skills from s…
-
GEPA method optimizes LLM prompts using AI critiques, no GPU needed
A new method called GEPA (Genetic-Pareto Evolutionary Prompt Adaptation) has been introduced, aiming to optimize LLM pipelines without requiring extensive GPU resources for fine-tuning. Developed by researchers from UC …
-
Robotics expert details reinforcement learning for robot simulation and training
Shawn Hymel has released a new video and is hosting a webinar on reinforcement learning for robotics. The video, part 2 of a series, covers Proximal Policy Optimization (PPO) and demonstrates training an agent to balanc…
-
New method Exp-RSFT enhances generative recommender systems
Researchers have introduced Exponential reward-weighted fine-tuning (Exp-RSFT), a novel method for improving generative recommender systems, particularly when dealing with sparse and noisy user feedback. This technique …
-
New framework REFINE-DP enhances humanoid robot loco-manipulation
Researchers have developed REFINE-DP, a novel hierarchical framework designed to enhance humanoid robot capabilities in loco-manipulation tasks. This approach integrates diffusion policies (DPs) with reinforcement learn…
-
New benchmark reveals limitations in LLM personalization
Researchers have introduced Personalized RewardBench, a new benchmark designed to evaluate how well reward models for large language models can capture individual user preferences. Existing state-of-the-art reward model…
-
New ReDiPPO framework boosts LLM mathematical reasoning
Researchers have introduced ReDiPPO, a novel framework designed to enhance the mathematical reasoning abilities of large language models. This approach addresses the challenge of accurate token-level credit assignment i…
-
Frozen CNNs in RL spontaneously develop sparse representations
Researchers have observed that deep reinforcement learning agents, when using frozen, randomly initialized Convolutional Neural Network (CNN) feature extractors, spontaneously develop highly sparse fully-connected repre…