PulseAugur
EN
LIVE 10:24:18

New research explores advanced RL for agent survival, navigation, and explainability · 7 sources tracked

Researchers are exploring advanced techniques in reinforcement learning (RL) to enhance agent performance and interpretability. One study introduces programmatic policies (PERL) as an alternative to neural policies (NERL) in evolutionary RL, demonstrating PERL's superior survival rates. Another paper focuses on flow-aware navigation for autonomous robots using RL, finding that agents trained with velocity and memory sensors outperform those given explicit global flow parameters. Additionally, new methods are being developed to explain RL agents, such as using Inductive Logic Programming (ILP) to create objective metrics for policy explainability, and a technique called SteinGate that improves safety by detecting rare catastrophic tail events in RL. AI

IMPACT These advancements in RL techniques could lead to more robust, efficient, and interpretable AI agents across various domains, from robotics to complex decision-making scenarios.

RANK_REASON Multiple arXiv papers published on distinct research topics within reinforcement learning.

Read on Hugging Face Daily Papers →

AI-generated summary · Google Gemini · from 304 sources. How we write summaries →

New research explores advanced RL for agent survival, navigation, and explainability · 7 sources tracked

COVERAGE [304]

  1. arXiv cs.AI TIER_1 English(EN) · Yuxuan Zhu, Rohan Alur, Daniel Kang ·

    Non-vacuous Generalization Bounds for Reinforcement Learning with Verifiable Rewards

    arXiv:2607.14506v1 Announce Type: cross Abstract: While reinforcement learning with verifiable rewards (RLVR) is widely used to improve the reasoning capabilities of large language models (LLMs), the generalizability of the resulting models remains poorly understood. In this work…

  2. arXiv cs.LG TIER_1 English(EN) · Asha Ramanujam (Davidson School of Chemical Engineering, Purdue University, West Lafayette, IN), Adam Elyoumi (Davidson School of Chemical Engineering, Purdue University, West Lafayette, IN), Hao Chen (Davidson School of Chemical Engineering, Purdue Univ… ·

    SafeOR-Gym: A Benchmark Suite for Safe Reinforcement Learning Algorithms on Practical Operations Research Problems

    arXiv:2506.02255v2 Announce Type: replace Abstract: Most existing safe reinforcement learning (RL) benchmarks focus on robotics and control tasks, offering limited relevance to high-stakes domains that involve structured constraints, mixed-integer decisions, and industrial comple…

  3. arXiv cs.LG TIER_1 English(EN) · Zikun Zhang, Jiayuan Sheng, David D. Yao, Wenpin Tang ·

    A Continuous-Time Reinforcement Learning Framework for Fine-Tuning Discrete Diffusion Models

    arXiv:2607.14522v1 Announce Type: new Abstract: We formulate reinforcement learning (RL) in continuous time with discrete state spaces and possibly arbitrary action spaces via a stochastic control approach, where the state dynamics are modeled as a controlled continuous-time Mark…

  4. arXiv cs.CL TIER_1 English(EN) · Bowei He, Yankai Chen, Xiaokun Zhang, Xue Liu ·

    Branching Policy Optimization: Sandbox-Native Language Agent Reinforcement Learning

    arXiv:2607.14171v1 Announce Type: cross Abstract: Reinforcement learning has emerged as the dominant paradigm for training large language model (LLM) agents that interact with executable sandboxes. State-of-the-art algorithms such as PPO, RLOO, and GRPO inherit their rollout topo…

  5. arXiv cs.CL TIER_1 English(EN) · Jinyang Wu, Shuo Yang, Zhengxi Lu, Fan Zhang, Yuhao Shen, Lang Feng, Haoran Luo, Zheng Lian, Shuai Zhang, Zhengqi Wen, Jianhua Tao ·

    SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning

    arXiv:2607.14777v1 Announce Type: new Abstract: Large language models are increasingly trained as interactive agents for long-horizon tasks involving multi-turn interaction, tool use, and environment feedback. Outcome-based reinforcement learning (RL) provides a practical optimiz…

  6. arXiv cs.AI TIER_1 English(EN) · Mattie Fellows, Clarisse Wibault, Uljad Berdica, Johannes Forkel, Maike Osborne, Jakob N. Foerster ·

    Fully Offline Reinforcement Learning

    arXiv:2505.22442v3 Announce Type: replace-cross Abstract: Offline RL (ORL) promises safe and sample-efficient deployment but existing methods rely on undocumented online interactions for hyperparameter tuning and lack reliable fully offline estimates of initial online performance…

  7. arXiv cs.AI TIER_1 English(EN) · Mohsen Amiri, Sindri Magn\'usson ·

    Reinforcement Learning in Switching Non-Stationary Markov Decision Processes: Algorithms and Convergence Analysis

    arXiv:2503.18607v2 Announce Type: replace-cross Abstract: We introduce the Switching Non-Stationary Markov Decision Process (SNS-MDP) framework, in which the environment transitions among a finite set of MDPs governed by a latent Markov chain while the agent observes only the ext…

  8. arXiv cs.CL TIER_1 English(EN) · Jianhua Tao ·

    SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning

    Large language models are increasingly trained as interactive agents for long-horizon tasks involving multi-turn interaction, tool use, and environment feedback. Outcome-based reinforcement learning (RL) provides a practical optimization paradigm, but its sparse trajectory-level …

  9. arXiv cs.LG TIER_1 English(EN) · Junyi Wu, Dan Li ·

    Factorized Spectral Representations for Reinforcement Learning

    arXiv:2607.13498v1 Announce Type: new Abstract: Learning a compact model of the world from interaction data is central to sample-efficient deep reinforcement learning. Spectral representation methods have become the leading paradigm for representation learning in continuous contr…

  10. arXiv cs.AI TIER_1 English(EN) · Celeste Veronese, Edoardo Zorzi, Daniele Meli, Alessandro Farinelli ·

    Explaining Reinforcement Learning Agents via Inductive Logic Programming

    arXiv:2607.13655v1 Announce Type: new Abstract: Explainable Reinforcement Learning (XRL) seeks to make Reinforcement Learning (RL) policies more transparent and interpretable, a key requirement in safety-critical and human-centric scenarios. However, it is mostly based on user st…

  11. arXiv cs.LG TIER_1 English(EN) · Andrea Maria Braghin, Nicol\`o Botteghi, Matteo Tomasetto, Andrea Manzoni, Gabriele Cazzulani ·

    Flow-aware Optimal Navigation in Unsteady Flows through Reinforcement Learning

    arXiv:2607.13553v1 Announce Type: cross Abstract: Autonomous robotic navigation in nonstationary time-varying fluid flows remains a fundamental challenge due to partial observability and the unpredictability of realistic environments. While classical optimal control frameworks em…

  12. arXiv cs.LG TIER_1 English(EN) · Merlin Paul, Anup Aprem ·

    Structured Reinforcement Learning for Bayesian Persuasion : Application to Intelligent Interactive Driving

    arXiv:2607.13576v1 Announce Type: new Abstract: Interactive driving, wherein an intelligent lead vehicle equipped with real-time traffic data coordinates route choices of connected vehicles, offers a promising approach to dynamic traffic management. To address the challenge of ha…

  13. arXiv cs.AI TIER_1 Deutsch(DE) · Yassine Chemingui, Chenhua Fan, Honghao Wei, Janardhan Rao Doppa ·

    SteinGate: Tail-Sensitive Safe Reinforcement Learning via Stein Discrepancy

    arXiv:2607.13175v1 Announce Type: cross Abstract: Safe reinforcement learning typically enforces safety by bounding expected cumulative costs, a criterion that often fails to detect rare but catastrophic tail events. To overcome these limitations, this paper introduces SteinGate,…

  14. arXiv cs.LG TIER_1 English(EN) · Jiale Fan, Andrei Cramariuc, Tifanny Portela, Marco Hutter ·

    Pretraining in Actor-Critic Reinforcement Learning for Locomotion

    arXiv:2510.12363v4 Announce Type: replace-cross Abstract: The pretraining-finetuning paradigm has facilitated numerous transformative advancements in artificial intelligence research in recent years. However, in the domain of reinforcement learning (RL) for robot locomotion, indi…

  15. arXiv cs.LG TIER_1 English(EN) · Anton Roupassov-Ruiz, Yiyang Zuo ·

    Survival Dynamics of Neural and Programmatic Policies in Evolutionary Reinforcement Learning

    arXiv:2601.04365v2 Announce Type: replace Abstract: In evolutionary reinforcement learning tasks (ERL), agent policies are often encoded as small artificial neural networks (NERL). Such representations lack explicit modular structure, limiting behavioral interpretation. We invest…

  16. arXiv cs.LG TIER_1 English(EN) · Wenpin Tang ·

    A Continuous-Time Reinforcement Learning Framework for Fine-Tuning Discrete Diffusion Models

    We formulate reinforcement learning (RL) in continuous time with discrete state spaces and possibly arbitrary action spaces via a stochastic control approach, where the state dynamics are modeled as a controlled continuous-time Markov chain (CTMC). We consider policy optimization…

  17. arXiv cs.LG TIER_1 English(EN) · Daniel Kang ·

    Non-vacuous Generalization Bounds for Reinforcement Learning with Verifiable Rewards

    While reinforcement learning with verifiable rewards (RLVR) is widely used to improve the reasoning capabilities of large language models (LLMs), the generalizability of the resulting models remains poorly understood. In this work, we establish the first non-vacuous generalizatio…

  18. Hugging Face Daily Papers TIER_1 English(EN) ·

    SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning

    Large language models are increasingly trained as interactive agents for long-horizon tasks involving multi-turn interaction, tool use, and environment feedback. Outcome-based reinforcement learning (RL) provides a practical optimization paradigm, but its sparse trajectory-level …

  19. arXiv cs.AI TIER_1 English(EN) · Alessandro Farinelli ·

    Explaining Reinforcement Learning Agents via Inductive Logic Programming

    Explainable Reinforcement Learning (XRL) seeks to make Reinforcement Learning (RL) policies more transparent and interpretable, a key requirement in safety-critical and human-centric scenarios. However, it is mostly based on user studies, thus targeting the needs of a specific au…

  20. arXiv cs.CL TIER_1 English(EN) · Xue Liu ·

    Branching Policy Optimization: Sandbox-Native Language Agent Reinforcement Learning

    Reinforcement learning has emerged as the dominant paradigm for training large language model (LLM) agents that interact with executable sandboxes. State-of-the-art algorithms such as PPO, RLOO, and GRPO inherit their rollout topology from RLHF: for each prompt, N independent tra…

  21. arXiv cs.LG TIER_1 English(EN) · Anup Aprem ·

    Structured Reinforcement Learning for Bayesian Persuasion : Application to Intelligent Interactive Driving

    Interactive driving, wherein an intelligent lead vehicle equipped with real-time traffic data coordinates route choices of connected vehicles, offers a promising approach to dynamic traffic management. To address the challenge of harmonising decisions, this paper considers the st…

  22. arXiv cs.LG TIER_1 English(EN) · Gabriele Cazzulani ·

    Flow-aware Optimal Navigation in Unsteady Flows through Reinforcement Learning

    Autonomous robotic navigation in nonstationary time-varying fluid flows remains a fundamental challenge due to partial observability and the unpredictability of realistic environments. While classical optimal control frameworks employed in robotics require unrealistic a-priori gl…

  23. arXiv cs.LG TIER_1 English(EN) · Dan Li ·

    Factorized Spectral Representations for Reinforcement Learning

    Learning a compact model of the world from interaction data is central to sample-efficient deep reinforcement learning. Spectral representation methods have become the leading paradigm for representation learning in continuous control by taking a matrix view of the transition ker…

  24. arXiv cs.AI TIER_1 English(EN) · A Run, Ziluo Ding ·

    In-Context Reinforcement Learning under Non-Stationarity: A Survey

    arXiv:2607.11906v1 Announce Type: new Abstract: The development of decision-pretrained transformers, algorithm distillation, long-context meta-RL, and retrieval-augmented agents has renewed interest in in-context reinforcement learning (ICRL): the ability of a pretrained or fine-…

  25. arXiv cs.LG TIER_1 English(EN) · Paolo Magliano, Puze Liu, Jan Peters, Davide Tateo, Raffaello Camoriano ·

    Directional Constraints for Efficient Exploration in Safe Reinforcement Learning

    arXiv:2607.12784v1 Announce Type: cross Abstract: Reinforcement Learning has revolutionized the landscape of robotic research, allowing robust learning of complex robotic skills in simulation. However, real-world deployment in open-ended environments requires strong safety guaran…

  26. arXiv cs.LG TIER_1 English(EN) · Amber Srivastava ·

    Environment Parameter Gradient Theorem for Policy-Environment Co-Design in Reinforcement Learning

    arXiv:2607.12590v1 Announce Type: cross Abstract: Reinforcement learning (RL) is traditionally concerned with learning a control policy for a fixed environment. In many engineering systems, however, the environment itself is alterable: physical or operational parameters can be tu…

  27. arXiv cs.AI TIER_1 English(EN) · Zhouchonghao Wu, Raymond Song, Vedant Mundheda, Luis E. Navarro-Serment, Christof Schoenborn, Jeff Schneider ·

    TADPO: Reinforcement Learning Goes Off-road

    arXiv:2603.05995v2 Announce Type: replace-cross Abstract: Off-road autonomous driving poses significant challenges such as navigating unmapped, variable terrain with uncertain and diverse dynamics. Addressing these challenges requires effective long-horizon planning and adaptable…

  28. arXiv cs.AI TIER_1 English(EN) · Tongxi Wang, Zhuoyang Xia, Xinran Chen, Shan Liu ·

    Tracking Drift: Variation-Aware Entropy Scheduling for Non-Stationary Reinforcement Learning

    arXiv:2601.19624v3 Announce Type: replace-cross Abstract: Real-world reinforcement learning often faces environment drift, but most existing methods rely on static entropy coefficients/target entropy, causing over-exploration during stable periods and under-exploration after drif…

  29. arXiv cs.AI TIER_1 English(EN) · Xin Qiu, Yulu Gan, Conor F. Hayes, Qiyao Liang, Yinggan Xu, Roberto Dailey, Elliot Meyerson, Babak Hodjat, Risto Miikkulainen ·

    Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning

    arXiv:2509.24372v3 Announce Type: replace-cross Abstract: Fine-tuning large language models (LLMs) for downstream tasks is an essential stage of modern AI deployment. Reinforcement learning (RL) has emerged as the dominant fine-tuning paradigm, underpinning many state-of-the-art …

  30. arXiv cs.AI TIER_1 English(EN) · Emil Mittag, Richard Dazeley, Peter Vamplew ·

    OOD-RL-Bench: A Benchmark Framework for Out-of-Distribution Detection in Reinforcement Learning

    arXiv:2607.12523v1 Announce Type: cross Abstract: Reliable reinforcement learning (RL) agents must maintain operational integrity amidst sensor malfunctions, dynamic disturbances, and slow environmental shifts. The detection of out-of-distribution conditions is pivotal to determi…

  31. arXiv cs.AI TIER_1 English(EN) · Jonas Ehrhardt, Ren\'e Heesch, Oliver Niggemann ·

    Knowledge- and Gradient-Guided Reinforcement Learning for Parametrized Action Markov Decision Processes

    arXiv:2607.12924v1 Announce Type: new Abstract: In this paper, we study Reinforcement Learning in Parametrized Action Markov Decision Processes (PAMDP), where each decision consists of a symbolic action and numerical parameters. In such settings Reinforcement Learning algorithms …

  32. arXiv cs.AI TIER_1 English(EN) · Yuhui Bie, Guowei Xu, Yaojun Wang ·

    Calibration-First Reward-Component Auditing for Reinforcement Learning Control in Smart Greenhouses

    arXiv:2607.11959v1 Announce Type: new Abstract: Greenhouse reinforcement learning can test climate-control ideas at a speed and scale that is difficult to achieve with crop experiments alone. For smart-greenhouse control, however, a single simulator return is not enough: a grower…

  33. arXiv cs.AI TIER_1 English(EN) · Oliver Niggemann ·

    Knowledge- and Gradient-Guided Reinforcement Learning for Parametrized Action Markov Decision Processes

    In this paper, we study Reinforcement Learning in Parametrized Action Markov Decision Processes (PAMDP), where each decision consists of a symbolic action and numerical parameters. In such settings Reinforcement Learning algorithms typically determine parameters with one-shot est…

  34. arXiv cs.LG TIER_1 English(EN) · Raffaello Camoriano ·

    Directional Constraints for Efficient Exploration in Safe Reinforcement Learning

    Reinforcement Learning has revolutionized the landscape of robotic research, allowing robust learning of complex robotic skills in simulation. However, real-world deployment in open-ended environments requires strong safety guarantees to prevent dangerous or harmful behaviors. Sa…

  35. Hugging Face Daily Papers TIER_1 English(EN) ·

    Directional Constraints for Efficient Exploration in Safe Reinforcement Learning

    Reinforcement Learning has revolutionized the landscape of robotic research, allowing robust learning of complex robotic skills in simulation. However, real-world deployment in open-ended environments requires strong safety guarantees to prevent dangerous or harmful behaviors. Sa…

  36. arXiv cs.LG TIER_1 English(EN) · Amber Srivastava ·

    Environment Parameter Gradient Theorem for Policy-Environment Co-Design in Reinforcement Learning

    Reinforcement learning (RL) is traditionally concerned with learning a control policy for a fixed environment. In many engineering systems, however, the environment itself is alterable: physical or operational parameters can be tuned to shape the transition dynamics and costs exp…

  37. arXiv cs.AI TIER_1 English(EN) · Peter Vamplew ·

    OOD-RL-Bench: A Benchmark Framework for Out-of-Distribution Detection in Reinforcement Learning

    Reliable reinforcement learning (RL) agents must maintain operational integrity amidst sensor malfunctions, dynamic disturbances, and slow environmental shifts. The detection of out-of-distribution conditions is pivotal to determining when an agent's observations, transitions, or…

  38. Hugging Face Daily Papers TIER_1 English(EN) ·

    OOD-RL-Bench: A Benchmark Framework for Out-of-Distribution Detection in Reinforcement Learning

    Reliable reinforcement learning (RL) agents must maintain operational integrity amidst sensor malfunctions, dynamic disturbances, and slow environmental shifts. The detection of out-of-distribution conditions is pivotal to determining when an agent's observations, transitions, or…

  39. arXiv cs.LG TIER_1 English(EN) · Jiaheng Hu, Zizhao Wang, Peter Stone, Roberto Mart\'in-Mart\'in ·

    Disentangled Unsupervised Skill Discovery for Efficient Hierarchical Reinforcement Learning

    arXiv:2410.11251v2 Announce Type: replace Abstract: A hallmark of intelligent agents is the ability to learn reusable skills purely from unsupervised interaction with the environment. However, existing unsupervised skill discovery methods often learn entangled skills where one sk…

  40. arXiv cs.LG TIER_1 English(EN) · Ekkachai Jueng ·

    Q-Learning Lab: Teaching Reinforcement Learning Through Learner-Generated Trace Analysis

    arXiv:2607.10802v1 Announce Type: cross Abstract: Reinforcement learning is usually introduced through the Bellman update, yet the equation often remains abstract to undergraduates: they watch policy arrows converge but rarely observe how each value is computed or why an action i…

  41. arXiv cs.LG TIER_1 English(EN) · Simone Drago, Marco Mussi, Leonardo Bianconi, Alberto Maria Metelli ·

    Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability

    arXiv:2607.11432v1 Announce Type: new Abstract: In this work, we study the reinforcement learning (RL) problem from pairwise trajectory comparisons provided by a human expert. We generalize preference-based RL by formalizing a novel setting in which the expert can also label traj…

  42. arXiv cs.LG TIER_1 English(EN) · Sai Anirudh Katupilla, Shreeya Dasa Lakshminath ·

    Nonparametric Bayesian Inverse Reinforcement Learning with Data-Parallel Gibbs Sampling

    arXiv:2607.09886v1 Announce Type: new Abstract: Inverse Reinforcement Learning recovers reward functions from expert demonstrations, but standard formulations assume that all demonstrations come from a single expert. When demonstrations are pooled from multiple experts with disti…

  43. arXiv cs.AI TIER_1 English(EN) · Bonan Wang, Letian Tao, Bin Shuai, Jiaxin Gao, Wenxin Zhao, Wei Xiong, Kehua Sheng, Bo Zhang, Yang Guan, Shengbo Eben Li ·

    FAST: A Framework for Aligned Sampling and Training in Parallel Reinforcement Learning for Autonomous Driving

    arXiv:2606.21587v2 Announce Type: replace-cross Abstract: Deep reinforcement learning is pivotal for closed-loop autonomous driving yet remains constrained by severe bottlenecks in sampling efficiency. Standard parallel sampling mitigates this but suffers from the straggler effec…

  44. arXiv cs.AI TIER_1 English(EN) · Xiaoya Li, Albert Wang, Guoyin Wang, Chris Shum, Jiwei Li ·

    CRINN: Contrastive Reinforcement Learning for Approximate Nearest Neighbor Search

    arXiv:2508.02091v3 Announce Type: replace-cross Abstract: Approximate nearest-neighbor search (ANNS) algorithms have become increasingly critical for recent AI applications, particularly in retrieval-augmented generation (RAG) and agent-based LLM applications. In this paper, we p…

  45. arXiv cs.AI TIER_1 English(EN) · John Wikman, Alexandre Proutiere, David Broman ·

    Adaptive Reinforcement Learning for Unobservable Random Delays

    arXiv:2506.14411v2 Announce Type: replace-cross Abstract: In standard reinforcement learning (RL) settings, the interaction between the agent and the environment is typically modeled as a Markov decision process (MDP), which assumes that the agent observes the system state instan…

  46. arXiv cs.AI TIER_1 English(EN) · Pengfei Cai, Utkarsh Utkarsh, Alan Edelman, Christopher Vincent Rackauckas, Rafael Gomez-Bombarelli ·

    Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards

    arXiv:2607.10474v1 Announce Type: cross Abstract: Partial differential equations (PDEs) are foundational to modeling in science and engineering, but constructing reliable numerical solvers remains labor-intensive, demanding expert knowledge of discretization schemes, stability co…

  47. arXiv cs.AI TIER_1 English(EN) · Wenke Xia, Pei Ren, Wenbo Yu, Yizhuo Zhang, Jifan Li, Yixue Zhang, Yinuo Zhao, Qingyang Gao, Jianlong Fu, Jian Tang, Ji-Rong Wen, Zhengping Che, Di Hu ·

    Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning

    arXiv:2607.09866v1 Announce Type: cross Abstract: Offline-to-online reinforcement learning is promising for generalizable robotic manipulation, yet its full-stack complexity obscures reproduction and diagnosis. Within such systems, value estimation plays a central role in priorit…

  48. arXiv cs.AI TIER_1 English(EN) · Roberto Garrone ·

    Transfer Learning Across Policy Regimes in Adaptive Multi-Agent Systems

    arXiv:2607.09685v1 Announce Type: cross Abstract: Policy models often assume that the relationship between a policy instrument and its outcome remains stable across institutional conditions. In adaptive socio-technical systems this assumption may fail: regulatory change can alter…

  49. arXiv cs.AI TIER_1 English(EN) · Mingyuan Wu, Jingcheng Yang, Shengyi Qian, Xudong Wang, Jize Jiang, Qifan Wang, Aashu Singh, Khoi Pham, Fei Liu, Zhaolun Su, Zhuokai Zhao, Klara Nahrstedt, Jianyu Wang, Hanchao Yu ·

    SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning

    arXiv:2607.10966v1 Announce Type: new Abstract: We introduce Self-Verified Reasoner (SVR-R1), a multi-turn RL framework that turns a model's own verification into a learning signal for multimodal reasoning. For each query, the model proposes an answer using the same weights, and …

  50. arXiv cs.AI TIER_1 English(EN) · Mianqiu Huang, Taofeng Xue, Chong Peng, Jinrui Ding, Sicheng Fan, Jiale Hong, Yufei Gao, Xiaocheng Zhang, Linsen Guo, Xin Yang, Dengchang Zhao, Yuchen Xie, Peng Pei, Xunliang Xie, Xipeng Qiu ·

    EvoCUA-1.5: Online Reinforcement Learning for Multi-turn Computer-Use Agents

    arXiv:2607.09773v1 Announce Type: new Abstract: Computer-use agents must solve long-horizon tasks through repeated interaction with partially observable, multimodal desktop environments. Although imitation learning and offline trajectory refinement provide strong priors, static t…

  51. arXiv cs.AI TIER_1 English(EN) · Rongping Zhou, Omid Tavallaie, Shuaijun Chen, Albert Y. Zomaya ·

    Measure the Sim-to-Real Gap: Designing an Affordable Real-World Benchmark Platform for Reinforcement Learning in AIoT Systems

    arXiv:2607.10309v1 Announce Type: new Abstract: Reinforcement learning (RL) is commonly employed to enhance the performance of autonomous systems, including the Autonomous Internet of Things (AIoT). However, the trial-and-error nature of RL, when conducted in real-world environme…

  52. arXiv cs.AI TIER_1 English(EN) · Rushuai Yang, Zhuo Han, Houlin Li, Hecheng Wang, Zhichao Wu, Rui Zhang, Zhaowei Zhang, Zihong Chen, Xiaohan Yan, Chiming Liu, Yi Chen, Wei Shan, Maoqing Yao ·

    VINE: Taming Generative Control Policies for Reinforcement Learning

    arXiv:2607.10369v1 Announce Type: cross Abstract: Flow-matching policies have emerged as an effective policy parameterization for robot learning. They iteratively generate actions from noise, enabling highly expressive modeling of complex and multimodal action distributions. Howe…

  53. arXiv cs.AI TIER_1 English(EN) · Baohe Zhang, Lilli Frison, Thomas Brox, Joschka B\"odecker ·

    Constrained Reinforcement Learning for Safe Heat Pump Control

    arXiv:2409.19716v2 Announce Type: replace-cross Abstract: Constrained Reinforcement Learning (RL) has emerged as a significant research area within RL, where integrating constraints with rewards is crucial for enhancing safety and performance across diverse control tasks. In the …

  54. arXiv cs.AI TIER_1 English(EN) · DeepReinforce Team, Xiaoya Li, Guoyin Wang, Songqiao Su, Chris Shum, Jiwei Li ·

    GrandCode: Achieving Grandmaster Level in Competitive Programming via Agentic Reinforcement Learning

    arXiv:2604.02721v2 Announce Type: replace Abstract: Competitive programming remains one of the last few human strongholds in coding against AI. The best AI system to date still underperforms the best humans competitive programming: the most recent best result, Google's Gemini~3 D…

  55. arXiv cs.AI TIER_1 English(EN) · Yunhai Feng, Natalie Leung, Jiaxuan Wang, Lujie Yang, Haozhi Qi, Preston Culbertson ·

    A Minimalist Retargeting-Guided Reinforcement Learning Recipe for Dexterous Manipulation

    arXiv:2607.11874v1 Announce Type: cross Abstract: Recent work in humanoid whole-body control has found success with a simple recipe: retarget human motion to robot kinematic references, then train policies via reinforcement learning (RL) to track them. But how does this recipe tr…

  56. arXiv cs.AI TIER_1 English(EN) · Alper Kamil Bozkurt, Shangtong Zhang, Yuichi Motai ·

    Active Offline-to-Online Reinforcement Learning

    arXiv:2607.11720v1 Announce Type: cross Abstract: Background: Offline reinforcement learning (RL) enables effective policies to be trained from large, previously collected datasets and subsequently improved through limited online interaction. This offline-to-online RL (O2O-RL) pa…

  57. arXiv cs.AI TIER_1 English(EN) · Preston Culbertson ·

    A Minimalist Retargeting-Guided Reinforcement Learning Recipe for Dexterous Manipulation

    Recent work in humanoid whole-body control has found success with a simple recipe: retarget human motion to robot kinematic references, then train policies via reinforcement learning (RL) to track them. But how does this recipe transfer to dexterous manipulation? The answer is no…

  58. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Alfred Höß ·

    Multi-Agent Reinforcement Learning for C-V2X RAT Selection

    Vehicles are increasingly equipped with advanced V2X communication capabilities. While early V2X apps utilized services such as Cooperative Awareness Messages, recent developments have allowed more advanced applications including cooperative driving, shared perception, and sensor…

  59. arXiv cs.AI TIER_1 English(EN) · Yuichi Motai ·

    Active Offline-to-Online Reinforcement Learning

    Background: Offline reinforcement learning (RL) enables effective policies to be trained from large, previously collected datasets and subsequently improved through limited online interaction. This offline-to-online RL (O2O-RL) paradigm is particularly promising in nonstationary …

  60. Hugging Face Daily Papers TIER_1 English(EN) ·

    Active Offline-to-Online Reinforcement Learning

    Background: Offline reinforcement learning (RL) enables effective policies to be trained from large, previously collected datasets and subsequently improved through limited online interaction. This offline-to-online RL (O2O-RL) paradigm is particularly promising in nonstationary …

  61. arXiv cs.LG TIER_1 English(EN) · Alberto Maria Metelli ·

    Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability

    In this work, we study the reinforcement learning (RL) problem from pairwise trajectory comparisons provided by a human expert. We generalize preference-based RL by formalizing a novel setting in which the expert can also label trajectory pairs as incomparable, i.e., when neither…

  62. arXiv cs.AI TIER_1 English(EN) · Jiayu Yao, Yiwei Wang, Anmeng Zhang, Zhe Sun, Songsong Wang, Lingrui Mei, Yuyao Ge, Shenghua Liu ·

    Multimodal Reward Hacking in Reinforcement Learning

    arXiv:2607.09492v1 Announce Type: new Abstract: Reinforcement learning (RL) is increasingly used to align multimodal large language models (MLLMs), but higher rewards do not always imply better task performance. This risk is amplified when visual evidence is evaluated by text-onl…

  63. arXiv cs.AI TIER_1 English(EN) · Guanquan Wang, Yoshimasa Tsuruoka ·

    Shortcut Trajectory Planning for Efficient Offline Reinforcement Learning

    arXiv:2607.09336v1 Announce Type: cross Abstract: Diffusion-based trajectory planners have shown strong performance in offline reinforcement learning, but their iterative denoising process often incurs high inference cost. Consistency-based planners reduce the number of sampling …

  64. arXiv cs.AI TIER_1 English(EN) · Brent Kong, Tejas Ram, Tony Yue Yu ·

    AlphaZero in Sparsely Rewarded Games: Limits and Auxiliary Supervision

    arXiv:2607.08984v1 Announce Type: cross Abstract: AlphaZero has demonstrated that a neural-guided Monte Carlo Tree Search can achieve superhuman performance, but strong play does not necessarily imply perfect play. We study this gap in two oracle-evaluable domains with contrastin…

  65. arXiv cs.LG TIER_1 English(EN) · Alan Nadelsticher Ruvalcaba ·

    Evolutionary Discovery of Developmental Reward Schedules in Deep Reinforcement Learning

    arXiv:2606.20858v2 Announce Type: replace Abstract: The temporal structure of reward composition in reinforcement learning (RL) is typically hand-designed and held fixed throughout training, leaving the progression of motivational priorities largely unexplored. In this work, we p…

  66. arXiv cs.LG TIER_1 English(EN) · Iris Xu, Sunshine Jiang, John Marangola, Nitish Dashora, Richard Li, Thomas Liu, Zexue He, Yuheng Zhi, Alex Pentland, Pulkit Agrawal, Zhang-Wei Hong ·

    Learning More from Less: Reinforcement Learning from Hindsight

    arXiv:2607.09042v1 Announce Type: new Abstract: Reinforcement learning (RL) is increasingly used to post-train vision-language-action (VLA) models, but every update consumes robot rollouts that are slow and costly to collect, making sample efficiency a central concern. Manipulati…

  67. arXiv cs.LG TIER_1 English(EN) · Elham Daneshmand, Majid Khadiv, Glen Berseth, Hsiu-Chin Lin ·

    SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions

    arXiv:2607.08925v1 Announce Type: new Abstract: Training reinforcement-learning agents directly on physical robots makes every fall costly, since a fall can damage the platform and cannot be undone like a simulator reset; the goal is therefore to minimize falls during training ra…

  68. arXiv cs.AI TIER_1 English(EN) · Tanmay Ambadkar, Sourav Panda, Shreyash Kale, Jonathan Dodge, Abhinav Verma ·

    Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization

    arXiv:2602.07764v2 Announce Type: replace-cross Abstract: Multi-objective reinforcement learning (MORL) seeks to train agents capable of balancing conflicting objectives. While single preference-conditioned policies offer a highly scalable solution, existing approaches remain bri…

  69. arXiv cs.AI TIER_1 English(EN) · Zihan Zhou, Animesh Garg, Ajay Mandlekar, Caelan Garrett ·

    ReinforceGen: Hybrid Skill Policies with Automated Data Generation and Reinforcement Learning

    arXiv:2512.16861v2 Announce Type: replace-cross Abstract: Long-horizon manipulation has been a long-standing challenge in the robotics community. We propose ReinforceGen, a system that combines task decomposition, data generation, imitation learning, and motion planning to form a…

  70. arXiv cs.AI TIER_1 English(EN) · Shenghua Liu ·

    Multimodal Reward Hacking in Reinforcement Learning

    Reinforcement learning (RL) is increasingly used to align multimodal large language models (MLLMs), but higher rewards do not always imply better task performance. This risk is amplified when visual evidence is evaluated by text-only or weakly grounded rewards. We study reward ha…

  71. arXiv cs.AI TIER_1 English(EN) · Yoshimasa Tsuruoka ·

    Shortcut Trajectory Planning for Efficient Offline Reinforcement Learning

    Diffusion-based trajectory planners have shown strong performance in offline reinforcement learning, but their iterative denoising process often incurs high inference cost. Consistency-based planners reduce the number of sampling steps, yet they typically rely on a two-stage teac…

  72. arXiv cs.AI TIER_1 English(EN) · Nivasini Ananthakrishnan, Mark Bedaywi, Michael I. Jordan, Stuart Russell, Nika Haghtalab ·

    Provably Optimal Learning Algorithms for Assistance Games

    arXiv:2607.08012v1 Announce Type: cross Abstract: This paper studies an online variant of the assistance games framework, where an informed agent and an uninformed agent repeatedly interact over $T$ timesteps to optimize a common reward function. While the informed agent (the hum…

  73. arXiv cs.AI TIER_1 English(EN) · Ezgi Korkmaz ·

    Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms

    arXiv:2607.07769v1 Announce Type: cross Abstract: Starting from the utilization of deep neural networks to approximate the state-action value function that led to winning one of the most challenging games, to algorithmic advancements that allowed solving problems without even exp…

  74. arXiv cs.AI TIER_1 English(EN) · Mumuksh Tayal, Manan Tayal, Ravi Prakash ·

    Safe Flow Q-Learning: Offline Safe Reinforcement Learning with Reachability-Based Flow Policies

    arXiv:2603.15136v2 Announce Type: replace-cross Abstract: Offline safe reinforcement learning (RL) seeks reward-maximizing policies from static datasets under strict safety constraints. Existing methods often rely on soft expected-cost objectives or iterative generative inference…

  75. arXiv cs.AI TIER_1 English(EN) · Xiuyi Lou, Zicheng Xu, Yu-Neng Chuang, Hoang Anh Duy Le, Zhaozhuo Xu, Guanchu Wang, Vladimir Braverman ·

    When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning

    arXiv:2607.07976v1 Announce Type: cross Abstract: Reinforcement learning (RL) has achieved remarkable success in enhancing the reasoning capabilities of large language models (LLMs). However, widely used critic-free RL methods rely on uniform credit assignment, broadcasting the s…

  76. arXiv cs.LG TIER_1 English(EN) · Nika Haghtalab, Chara Podimata, Kunhe Yang ·

    Calibrated Stackelberg Games: Learning Optimal Commitments Against Calibrated Agents

    arXiv:2306.02704v2 Announce Type: replace-cross Abstract: We introduce \emph{Calibrated Stackelberg Games (CSGs)}, a generalization of the standard Stackelberg Games (SGs) framework. In CSGs, a principal repeatedly interacts with an agent who (contrary to standard SGs) does not h…

  77. arXiv cs.CL TIER_1 English(EN) · Weiqi Wang, Xin Liu, Binxuan Huang, Hejie Cui, Rongzhi Zhang, Changlong Yu, Shuowei Jin, Jingfeng Yang, Qingyu Yin, Zhengyang Wang, Zheng Li, Yifan Gao, Priyanka Nigam, Bing Yin, Lihong Li, Yangqiu Song ·

    HeaPA: Difficulty-Aware Heap Sampling and On-Policy Query Augmentation for LLM Reinforcement Learning

    arXiv:2601.22448v2 Announce Type: replace-cross Abstract: RLVR has become a standard recipe for training LLMs on reasoning tasks with verifiable outcomes, but when rollout generation dominates the cost, efficiency hinges on which prompts are sampled and when. In practice, prompt …

  78. arXiv cs.LG TIER_1 English(EN) · Zhang-Wei Hong ·

    Learning More from Less: Reinforcement Learning from Hindsight

    Reinforcement learning (RL) is increasingly used to post-train vision-language-action (VLA) models, but every update consumes robot rollouts that are slow and costly to collect, making sample efficiency a central concern. Manipulation tasks typically provide only sparse rewards, …

  79. arXiv cs.LG TIER_1 English(EN) · Tony Yue Yu ·

    AlphaZero in Sparsely Rewarded Games: Limits and Auxiliary Supervision

    AlphaZero has demonstrated that a neural-guided Monte Carlo Tree Search can achieve superhuman performance, but strong play does not necessarily imply perfect play. We study this gap in two oracle-evaluable domains with contrasting structure: Connect Four, a solved partisan game …

  80. arXiv cs.AI TIER_1 English(EN) · Zhenyu Hou, Yujiang Li, Jie Tang, Yuxiao Dong ·

    Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

    arXiv:2607.07508v1 Announce Type: cross Abstract: Reinforcement learning (RL) is becoming increasingly important for post-training large language models (LLMs). Previous RL pipelines for LLMs were mostly synchronous and batch-interleaved, which is inefficient for long-horizon age…

  81. arXiv cs.LG TIER_1 English(EN) · Hokyun Im, Andrey Kolobov, Jianlong Fu, Youngwoon Lee ·

    Latent Policy Steering through One-Step Flow Policies

    arXiv:2603.05296v2 Announce Type: replace-cross Abstract: Offline reinforcement learning (RL) allows robots to learn from offline datasets without risky exploration. Yet, offline RL's performance often hinges on a brittle trade-off between (1) return maximization, which can push …

  82. arXiv cs.LG TIER_1 English(EN) · Denis Belomestny, Alexander Gasnikov, Egor Gladin, Alexey Naumov, Artemy Rubtsov, Yuri Sapronov, Daniil Tiapkin, Nikita Yudin ·

    Mathematical methods of reinforcement learning

    arXiv:2607.06935v1 Announce Type: cross Abstract: Reinforcement learning (RL) is increasingly grounded in tools from probability, optimization, and operator theory. This survey organizes the mathematical structures that underpin the design and analysis of modern algorithms in RL.…

  83. arXiv cs.LG TIER_1 English(EN) · Georg Sch\"afer, Jakob Rehrl, Stefan Huber, Simon Hirlaender ·

    Safe Reinforcement Learning using Ideas from Model Predictive Control

    arXiv:2607.07252v1 Announce Type: new Abstract: Reinforcement learning (RL) enables the synthesis of control policies directly from data, making it highly appealing for complex cyber-physical systems (CPSs) and robotics. A persistent challenge, however, is ensuring strict, hard s…

  84. arXiv cs.AI TIER_1 English(EN) · Ignacio D. Lopez-Miguel, Ezio Bartocci, Thomas Eiter, Martin Tappler ·

    ORCAID: Oblique Rule-Based Continuous-Action Interpretation for Deep RL Policies

    arXiv:2607.07235v1 Announce Type: cross Abstract: Explainability remains a key issue in reinforcement learning (RL). Distilling an interpretable policy from an agent trained in a complex environment is particularly challenging when the action space is continuous. We introduce ORC…

  85. arXiv cs.AI TIER_1 English(EN) · Zetian Hu, Shunyu Liu, Junjie Zhang, Yongcheng Jing, Ting-En Lin, Yongbin Li, Dacheng Tao ·

    Entropy Pacing Policy Optimization for Multi-Task Agentic Reinforcement Learning

    arXiv:2607.07178v1 Announce Type: cross Abstract: Recent breakthroughs of Reinforcement Learning (RL) have highlighted its potential for complex agentic Large Language Model (LLM) tasks. However, existing efforts largely focus on single-task settings, whereas real-world deploymen…

  86. arXiv cs.AI TIER_1 English(EN) · Dennis Gross, Quentin Mazouni, Helge Spieker, Arnaud Gotlieb ·

    Gimitest: A Comprehensive Tool for Testing Reinforcement Learning Policies

    arXiv:2607.07029v1 Announce Type: cross Abstract: Reinforcement learning (RL) policies can be unsafe and vulnerable to attacks. Ensuring their reliability is often a pain point as existing automated testing methods target only selected environments, testing scenarios, and RL algo…

  87. arXiv cs.CL TIER_1 English(EN) · Vladimir Braverman ·

    When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning

    Reinforcement learning (RL) has achieved remarkable success in enhancing the reasoning capabilities of large language models (LLMs). However, widely used critic-free RL methods rely on uniform credit assignment, broadcasting the same advantage to all tokens regardless of their di…

  88. Hugging Face Daily Papers TIER_1 English(EN) ·

    Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

    Reinforcement learning (RL) is becoming increasingly important for post-training large language models (LLMs). Previous RL pipelines for LLMs were mostly synchronous and batch-interleaved, which is inefficient for long-horizon agentic tasks. Recently, asynchronous RL has emerged …

  89. arXiv cs.AI TIER_1 English(EN) · Yuxiao Dong ·

    Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

    Reinforcement learning (RL) is becoming increasingly important for post-training large language models (LLMs). Previous RL pipelines for LLMs were mostly synchronous and batch-interleaved, which is inefficient for long-horizon agentic tasks. Recently, asynchronous RL has emerged …

  90. Hugging Face Daily Papers TIER_1 English(EN) ·

    Safe Reinforcement Learning using Ideas from Model Predictive Control

    Reinforcement learning (RL) enables the synthesis of control policies directly from data, making it highly appealing for complex cyber-physical systems (CPSs) and robotics. A persistent challenge, however, is ensuring strict, hard safety constraints during the active learning pha…

  91. arXiv cs.LG TIER_1 English(EN) · Simon Hirlaender ·

    Safe Reinforcement Learning using Ideas from Model Predictive Control

    Reinforcement learning (RL) enables the synthesis of control policies directly from data, making it highly appealing for complex cyber-physical systems (CPSs) and robotics. A persistent challenge, however, is ensuring strict, hard safety constraints during the active learning pha…

  92. arXiv cs.AI TIER_1 English(EN) · Martin Tappler ·

    ORCAID: Oblique Rule-Based Continuous-Action Interpretation for Deep RL Policies

    Explainability remains a key issue in reinforcement learning (RL). Distilling an interpretable policy from an agent trained in a complex environment is particularly challenging when the action space is continuous. We introduce ORCAID, a novel method for extracting interpretable r…

  93. Hugging Face Daily Papers TIER_1 English(EN) ·

    Entropy Pacing Policy Optimization for Multi-Task Agentic Reinforcement Learning

    Recent breakthroughs of Reinforcement Learning (RL) have highlighted its potential for complex agentic Large Language Model (LLM) tasks. However, existing efforts largely focus on single-task settings, whereas real-world deployment necessitates a generalist agent capable of solvi…

  94. arXiv cs.AI TIER_1 English(EN) · Dacheng Tao ·

    Entropy Pacing Policy Optimization for Multi-Task Agentic Reinforcement Learning

    Recent breakthroughs of Reinforcement Learning (RL) have highlighted its potential for complex agentic Large Language Model (LLM) tasks. However, existing efforts largely focus on single-task settings, whereas real-world deployment necessitates a generalist agent capable of solvi…

  95. arXiv cs.AI TIER_1 English(EN) · Arnaud Gotlieb ·

    Gimitest: A Comprehensive Tool for Testing Reinforcement Learning Policies

    Reinforcement learning (RL) policies can be unsafe and vulnerable to attacks. Ensuring their reliability is often a pain point as existing automated testing methods target only selected environments, testing scenarios, and RL algorithms. To address this, we propose a comprehensiv…

  96. arXiv cs.AI TIER_1 English(EN) · Muhammad Zain Amin, Kibele Sebnem Yildirim ·

    Self-Review Reinforcement Learning (SRRL) with Cross-Episode Memory and Policy Distillation

    arXiv:2607.05541v1 Announce Type: cross Abstract: Reinforcement Learning is commonly used to train large language models using environmental feedback. In applied settings, the environment usually provides sparse or delayed feedback. This makes it difficult for the model to pinpoi…

  97. arXiv cs.AI TIER_1 English(EN) · Yuhang Zhou, Kai Zheng, Haoling Li, Dengyun Peng, Can Xu, Jingjing Chen ·

    TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training

    arXiv:2607.05804v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student policy by matching a stronger teacher on the student's own trajectories, offering a promising framework for language agent training. However, its application to long-horizon agentic task…

  98. arXiv cs.AI TIER_1 English(EN) · Akshay Arora, Ishan Nigam, Ashutosh Aggarwal, Shefali Bansal, Krishna Singh, Sweta Kumari, Nikhil Mittal, Shariq Farhan, Siddarth Malreddy ·

    Beyond Static Evaluation: Building Simulation Environments for Scalable Agentic Reinforcement Learning

    arXiv:2607.05773v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve into autonomous agents, traditional static evaluation fails to capture multi-step decision-making. We introduce AgenticAI-Supervisor, an API and UI-driven RL Gym environment that decouples envi…

  99. arXiv cs.AI TIER_1 English(EN) · Haiwen Yi, Xinyuan Song ·

    Learning to Control LLM Agent Harnesses with Offline Reinforcement Learning

    arXiv:2607.05458v1 Announce Type: cross Abstract: Large language model (LLM) agents are usually improved by changing prompts, models, or hand-written workflows, while the execution harness around the model is treated as fixed infrastructure. We argue that this harness is itself a…

  100. arXiv cs.LG TIER_1 English(EN) · Gojko Perovic, Nuno Ferreira Duarte, Atabak Dehban, Gon\c{c}alo Teixeira, Egidio Falotico, Jos\'e Santos-Victor ·

    HERB: Human-augmented Efficient Reinforcement learning for Bin-packing

    arXiv:2504.16595v2 Announce Type: replace-cross Abstract: Packing objects efficiently is a fundamental problem in logistics, warehouse automation, and robotics. When dealing with highly diverse 3D objects (household or grocery items), closed-form solutions are infeasible, and heu…

  101. arXiv cs.LG TIER_1 English(EN) · Shuze Daniel Liu, Claire Chen, Jiabao Sean Xiao, Xin Chen, David Simchi-Levi ·

    Strategic Bargaining in Multi-Buyer Markets: Reinforcement Learning from Verifiable Rewards for LLM Negotiations

    arXiv:2607.05863v1 Announce Type: new Abstract: Negotiation is a fundamental strategic interaction in management science, characterized by agents attempting to reach agreements while protecting private information, such as reservation costs and hidden valuations. A prevalent yet …

  102. arXiv cs.AI TIER_1 English(EN) · Alexander Rombach, Chantale Lauer, Nijat Mehdiyev ·

    Improving LLM-Generated Process Model Quality Through Reinforcement Learning: The Role of Reward Function Design

    arXiv:2607.06175v1 Announce Type: cross Abstract: Large language models (LLMs) can generate BPMN process models from natural-language descriptions, yet supervised fine-tuning (SFT) limits their output quality to the patterns present in the training data. Reinforcement learning (R…

  103. arXiv cs.AI TIER_1 English(EN) · Yijun Zhang, Fan Xu, Jiaxin Ding, Yule Xie, Shiqing Gao, Xin Ding, Haoxiang Zhang, Luoyi Fu, Xinbing Wang ·

    Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents

    arXiv:2607.06223v1 Announce Type: new Abstract: Reinforcement learning has become a promising paradigm for improving large language model (LLM) agents on long-horizon search tasks, where the agent must make a sequence of intermediate decisions before receiving a final outcome. Ho…

  104. Hugging Face Daily Papers TIER_1 English(EN) ·

    Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

    Asynchronous reinforcement learning with single-rollout optimization addresses stability issues in LLM training for complex tasks, outperforming existing methods in coding and reasoning benchmarks.

  105. Hugging Face Daily Papers TIER_1 English(EN) ·

    Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms

    Starting from the utilization of deep neural networks to approximate the state-action value function that led to winning one of the most challenging games, to algorithmic advancements that allowed solving problems without even explicitly stating the rules of the challenge at hand…

  106. arXiv cs.AI TIER_1 English(EN) · Xinbing Wang ·

    Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents

    Reinforcement learning has become a promising paradigm for improving large language model (LLM) agents on long-horizon search tasks, where the agent must make a sequence of intermediate decisions before receiving a final outcome. However, existing methods still face a key limitat…

  107. arXiv cs.AI TIER_1 English(EN) · Nijat Mehdiyev ·

    Improving LLM-Generated Process Model Quality Through Reinforcement Learning: The Role of Reward Function Design

    Large language models (LLMs) can generate BPMN process models from natural-language descriptions, yet supervised fine-tuning (SFT) limits their output quality to the patterns present in the training data. Reinforcement learning (RL) can optimize beyond this ceiling using external…

  108. arXiv cs.LG TIER_1 English(EN) · David Simchi-Levi ·

    Strategic Bargaining in Multi-Buyer Markets: Reinforcement Learning from Verifiable Rewards for LLM Negotiations

    Negotiation is a fundamental strategic interaction in management science, characterized by agents attempting to reach agreements while protecting private information, such as reservation costs and hidden valuations. A prevalent yet complex scenario involves a single seller negoti…

  109. arXiv cs.AI TIER_1 English(EN) · Faid Keddouri, Sohaib Houhou, Aissa Boulmerka, Nadir Farhi ·

    Regime-Conditional Stabilisation of LLM-Augmented Cooperative Multi-Agent Reinforcement Learning

    arXiv:2607.04470v1 Announce Type: cross Abstract: Large Language Models (LLMs) offer a natural interface for translating human objectives into reward signals for cooperative multi-agent reinforcement learning (MARL), yet the training-time dynamics of this integration remain poorl…

  110. arXiv cs.AI TIER_1 English(EN) · Weiyang Guo, Zesheng Shi, Longhui Zhang, Zeen Zhu, Min Zhang, Jing Li ·

    Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry

    arXiv:2607.03702v1 Announce Type: new Abstract: Large language model (LLM) agents have shown strong decision-making capabilities in long-horizon interactive tasks, yet they still struggle to effectively leverage failed trajectories: full retries incur high interaction costs, whil…

  111. arXiv cs.AI TIER_1 English(EN) · Mingxuan Fan, Peiyang Liu ·

    Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning

    arXiv:2607.04242v1 Announce Type: new Abstract: Group-based reinforcement learning (RL) has become an effective paradigm for improving large language model agents on long-horizon interactive tasks. To obtain finer-grained policy updates than trajectory-level optimization, recent …

  112. arXiv cs.AI TIER_1 English(EN) · Qiuyi Qi, Jinjian Zhang, Mutian Bao, Tian Liang, Guocong Li, Dongnan Liu, Wei Zhou, Jie Liu, Ming Kong, Linjian Mo, Feng Zhang, Qiang Zhu ·

    CARL: Constraint-Aware Reinforcement Learning for Planning with LLMs

    arXiv:2607.04854v1 Announce Type: new Abstract: Despite their strong reasoning capabilities and extensive world knowledge, Large Language Models (LLMs) frequently generate plans that violate task constraints, undermining their reliability in real-world applications. This deficien…

  113. arXiv cs.AI TIER_1 English(EN) · Kai Zhao ·

    Mask-based Predictive Representations for Reinforcement Learning

    arXiv:2607.04153v1 Announce Type: cross Abstract: Vision-based deep reinforcement learning involves dealing with high-dimensional inputs of image information. It is crucial to abstract effective states from high-dimensional image inputs and limited samples for sample-efficient re…

  114. arXiv cs.AI TIER_1 English(EN) · Angen Ye, Weijie Ke, Xiaofeng Wang, Xinze Chen, Chaojun Ni, Guosheng Zhao, Boyuan Wang, Zheng Zhu, Junjie Xie, Dapeng Zhang ·

    HALO-WA: Hybrid-Attention Latent-Guided Online Reinforcement Learning for World-Action Models

    arXiv:2607.04265v1 Announce Type: cross Abstract: World-action (WA) models can generate long-horizon action chunks for general-purpose robotic manipulation, but they remain vulnerable to calibration, perception, and contact-dynamics errors in real-world precision tasks, often fai…

  115. arXiv cs.AI TIER_1 English(EN) · Qiang Liu, Taian Guo, Ruizhi Qiao, Xing Sun ·

    RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents

    arXiv:2607.04713v1 Announce Type: cross Abstract: Reinforcement learning holds significant potential for training large language models (LLMs) to handle multi-turn interactive tasks. However, in long-horizon, multi-turn tasks characterized by sparse outcome rewards, directly trai…

  116. arXiv cs.AI TIER_1 English(EN) · Saksham Sahai Srivastava, Vaneet Aggarwal ·

    A Technical Survey of Reinforcement Learning Techniques for Large Language Models

    arXiv:2507.04136v2 Announce Type: replace Abstract: This survey offers a comprehensive foundation on the integration of RL with language models, highlighting prominent algorithms such as Proximal Policy Optimization (PPO), Q-Learning, and Actor-Critic methods. Additionally, it pr…

  117. arXiv cs.AI TIER_1 English(EN) · Siqi Zhu, Jiaxuan You ·

    OpenTinker: Separating Concerns in Agentic Reinforcement Learning

    arXiv:2601.07376v2 Announce Type: replace Abstract: We introduce \textsc{OpenTinker}, an open infrastructure for training large language model (LLM) agents with many LoRA-backed policies over shared execution resources. Modern agent workloads mix supervised fine-tuning (SFT), onl…

  118. arXiv cs.AI TIER_1 English(EN) · Xiaoxuan Wang, Han Zhang, Haixin Wang, Yidan Shi, Ruoyan Li, Kaiqiao Han, Chenyi Tong, Haoran Deng, Renliang Sun, Alexander Taylor, Yanqiao Zhu, Jason Cong, Yizhou Sun, Wei Wang ·

    ARLArena: A Unified Framework for Stable Agentic Reinforcement Learning

    arXiv:2602.21534v3 Announce Type: replace Abstract: Agentic reinforcement learning (ARL) has rapidly gained attention as a promising paradigm for training agents to solve complex, multi-step interactive tasks. Despite encouraging early results, ARL remains highly unstable, often …

  119. arXiv cs.AI TIER_1 English(EN) · David Schiff, Ofir Lindenbaum, Yonathan Efroni ·

    ICR-RL: Deep Reinforcement Learning via In-Context Regression

    arXiv:2509.11259v2 Announce Type: replace-cross Abstract: Recent advancements in machine learning have largely been driven by foundation models (FMs) trained on large, diverse datasets, enabling them to generalize effectively to new, related tasks. However, extending this paradig…

  120. arXiv cs.AI TIER_1 English(EN) · Lu Li, Tianwei Ni, Yihao Sun, Pierre-Luc Bacon ·

    The Three Regimes of Offline-to-Online Reinforcement Learning

    arXiv:2510.01460v4 Announce Type: replace-cross Abstract: Offline-to-online reinforcement learning (RL) has emerged as a practical paradigm that leverages offline datasets for pretraining and online interactions for fine-tuning. However, its empirical behavior is highly inconsist…

  121. arXiv cs.AI TIER_1 English(EN) · Johannes Ackermann, Michael Noukhovitch, Takashi Ishida, Masashi Sugiyama ·

    Gradient Regularization Mitigates Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards

    arXiv:2602.18037v2 Announce Type: replace-cross Abstract: Reinforcement Learning from Human Feedback (RLHF) or Verifiable Rewards (RLVR) are two key steps in the post-training of modern Language Models (LMs). A common problem is reward hacking, where the policy may exploit inaccu…

  122. arXiv cs.AI TIER_1 English(EN) · Siyuan Wang, Lei Lei, Pranav Maheshwari, Sam Bellefeuille, Kan Zheng ·

    Multi-Agent Reinforcement Learning for V2X Resource Allocation: Disentangling MARL Challenges Through Benchmarking

    arXiv:2603.06607v2 Announce Type: replace-cross Abstract: Radio resource allocation (RRA) is a critical function in cellular vehicle-to-everything (C-V2X) networks, where vehicles must share limited wireless resources to support safety-critical communications. Multi-agent reinfor…

  123. arXiv cs.CL TIER_1 English(EN) · Mickel Liu, Liwei Jiang, Yancheng Liang, Simon Shaolei Du, Yejin Choi, Tim Althoff, Natasha Jaques ·

    Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models

    arXiv:2506.07468v4 Announce Type: replace-cross Abstract: Conventional large language model (LLM) safety alignment relies on a reactive, disjoint loop: attackers exploit a static model, then defenders patch exposed vulnerabilities. This sequential setup leads to attackers overfit…

  124. arXiv cs.LG TIER_1 English(EN) · Georg Sch\"afer, Jakob Rehrl, Stefan Huber, Simon Hirlaender ·

    Sample-Efficient Pareto Front Modeling for Energy-Aware Reinforcement Learning Using Bayesian Optimization

    arXiv:2607.03140v1 Announce Type: new Abstract: Industrial automation increasingly demands control strategies that balance operational performance with strict energy efficiency requirements. A common approach to solving this multi-objective problem, particularly within the framew…

  125. arXiv cs.LG TIER_1 English(EN) · Jiayi Guan, Tianle Zhang, Li Shen, Ruiqi Zhang, Ao Zhou, Lusong Li, Guai Chen, Mengjie Li, Alois Knoll, Xiaodong He, Changjun Jiang ·

    CDCP: Conditional Diffusion Model with Contextual Prompts for Multi-task Offline Safe Reinforcement Learning

    arXiv:2607.03903v1 Announce Type: new Abstract: Multi-task offline safe reinforcement learning (RL) promises to learn a shared optimal safe policy from offline data across multiple tasks. This paradigm provides an effective means for the widespread application of RL in multi-task…

  126. arXiv cs.LG TIER_1 English(EN) · Chee Heng Tan, Zhuoyi Lin, Mehul Motani, Wee Sun Lee ·

    On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models

    arXiv:2607.04332v1 Announce Type: new Abstract: In this paper, we consider the setting where large language models (LLMs) are trained using reinforcement learning (RL) to simultaneously improve reasoning accuracy and verbalize its confidence. Our reward scheme uses two functions …

  127. arXiv cs.LG TIER_1 English(EN) · Mao-Lin Luo, Zhe-Xu Wang, Zi-Hao Zhou, Bo Ye, Jian Zhao, Min-Ling Zhang, Tong Wei ·

    RL Forgets! Towards Continual Policy Optimization

    arXiv:2607.04364v1 Announce Type: new Abstract: Continual post-training is becoming a central paradigm for adapting vision-language models to evolving tasks. Recent work has increasingly favored reinforcement learning over supervised fine-tuning, driven by the belief that reinfor…

  128. arXiv cs.LG TIER_1 English(EN) · Fan Feng, Yujia Zheng, Minghao Fu, Yongqiang Chen, Guangyi Chen, Kevin Murphy, Biwei Huang, Kun Zhang ·

    Learning Task-Sufficient World Models by Synergizing Agentic Exploration and Structured Modeling

    arXiv:2607.04409v1 Announce Type: new Abstract: Learning and planning in imagination using world models provides an effective paradigm for training agents for decision-making. However, existing approaches often rely on high-dimensional latent spaces or generic visual embeddings t…

  129. arXiv cs.LG TIER_1 English(EN) · Kyohei Suzuki, onstantinos Slavakis ·

    Non-Convex Sparse Reinforcement Learning via Non-Monotone Inclusions

    arXiv:2607.04990v1 Announce Type: new Abstract: This work delivers two key contributions: one to efficient feature selection in reinforcement learning (RL), the other to the theory of non-monotone inclusions. On the RL side, the estimation bias inherent in conventional regulariza…

  130. arXiv cs.LG TIER_1 English(EN) · Yujiang Li, Zhenyu Hou, Yi Jing, Jie Tang, Yuxiao Dong ·

    CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents

    arXiv:2607.05378v1 Announce Type: new Abstract: Long-horizon agentic LLMs are increasingly limited by finite context windows, as extended interaction trajectories can exceed the maximum context length before a task is completed. Context compaction offers a natural solution by sum…

  131. arXiv cs.LG TIER_1 English(EN) · Jialun Cao, Fernando Acero, David \v{S}i\v{s}ka, Yufei Zhang ·

    Entropy Regularization Improves Policy Robustness in Continuous-Time Reinforcement Learning

    arXiv:2607.03168v1 Announce Type: cross Abstract: Entropy regularization is widely used in continuous-time reinforcement learning (RL) to reduce sensitivity to environmental perturbations, yet its robustness benefits lack a rigorous theoretical foundation. This paper establishes …

  132. arXiv cs.LG TIER_1 English(EN) · Manuel Wendl, Lukas Koller, Tobias Ladner, Matthias Althoff ·

    Training Verifiably Robust Agents Using Set-Based Reinforcement Learning

    arXiv:2408.09112v2 Announce Type: replace Abstract: Reinforcement learning policies parametrized by deep neural networks have achieved strong performance for continuous control, yet even small input perturbations may lead to unpredictable behavior. This sensitivity limits their u…

  133. arXiv cs.LG TIER_1 English(EN) · Waris Radji, Thomas Michel, Hector Piteau ·

    Octax: Accelerated CHIP-8 Arcade Environments for Reinforcement Learning in JAX

    arXiv:2510.01764v3 Announce Type: replace Abstract: Reinforcement learning (RL) research requires diverse, challenging environments that are both tractable and scalable. While modern video games may offer rich dynamics, they are computationally expensive and poorly suited for lar…

  134. arXiv cs.LG TIER_1 English(EN) · Vidur Sinha, Muhammed Ustaomeroglu, Guannan Qu ·

    Transformer-Based Multi-Agent Reinforcement Learning for Networked Systems with Long-Range Interactions

    arXiv:2511.13103v2 Announce Type: replace Abstract: Multi-agent reinforcement learning (MARL) has shown promise for large-scale network control, yet existing methods face two major limitations. First, they typically rely on an exponential decay property of agent interactions on f…

  135. arXiv cs.LG TIER_1 English(EN) · Tianshi Xu, Yuteng Chen, Meng Li ·

    CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning

    arXiv:2601.15141v2 Announce Type: replace Abstract: Agentic Reinforcement Learning (RL) has empowered Large Language Models (LLMs) to utilize tools like Python interpreters for complex problem-solving. However, for parameter-constrained models (e.g., 4B--7B), the exploration phas…

  136. arXiv cs.LG TIER_1 English(EN) · Bang Giang Le, Viet Cuong Ta ·

    Explicit Credit Assignment through Local Rewards and Dependence Graphs in Multi-Agent Reinforcement Learning

    arXiv:2601.21523v2 Announce Type: replace Abstract: To promote cooperation in Multi-Agent Reinforcement Learning, the reward signals of all agents can be aggregated together, forming global rewards that are commonly known as the fully cooperative setting. However, global rewards …

  137. arXiv cs.CL TIER_1 English(EN) · Jingjing Chen ·

    TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training

    On-policy distillation (OPD) trains a student policy by matching a stronger teacher on the student's own trajectories, offering a promising framework for language agent training. However, its application to long-horizon agentic tasks remains insufficiently explored. We identify t…

  138. Hugging Face Daily Papers TIER_1 English(EN) ·

    TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training

    Turn-level budgeting strategy for efficient on-policy distillation in long-horizon agent training addresses inefficiencies in full-horizon rollouts and shallow token concentration.

  139. Hugging Face Daily Papers TIER_1 English(EN) ·

    RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures

    RoboTALES introduces a two-stage framework that combines LLM-based planning and VLM-based criticism to improve task-aligned video generation and robotic policy training.

  140. arXiv cs.LG TIER_1 English(EN) · Yuxiao Dong ·

    CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents

    Long-horizon agentic LLMs are increasingly limited by finite context windows, as extended interaction trajectories can exceed the maximum context length before a task is completed. Context compaction offers a natural solution by summarizing previous interaction states and continu…

  141. arXiv cs.LG TIER_1 English(EN) · onstantinos Slavakis ·

    Non-Convex Sparse Reinforcement Learning via Non-Monotone Inclusions

    This work delivers two key contributions: one to efficient feature selection in reinforcement learning (RL), the other to the theory of non-monotone inclusions. On the RL side, the estimation bias inherent in conventional regularization schemes is addressed by augmenting classica…

  142. Hugging Face Daily Papers TIER_1 English(EN) ·

    CARL: Constraint-Aware Reinforcement Learning for Planning with LLMs

    Despite their strong reasoning capabilities and extensive world knowledge, Large Language Models (LLMs) frequently generate plans that violate task constraints, undermining their reliability in real-world applications. This deficiency arises from a lack of systematic mechanisms t…

  143. arXiv cs.AI TIER_1 English(EN) · Qiang Zhu ·

    CARL: Constraint-Aware Reinforcement Learning for Planning with LLMs

    Despite their strong reasoning capabilities and extensive world knowledge, Large Language Models (LLMs) frequently generate plans that violate task constraints, undermining their reliability in real-world applications. This deficiency arises from a lack of systematic mechanisms t…

  144. Hugging Face Daily Papers TIER_1 English(EN) ·

    Learning to Control LLM Agent Harnesses with Offline Reinforcement Learning

    Large language model (LLM) agents are usually improved by changing prompts, models, or hand-written workflows, while the execution harness around the model is treated as fixed infrastructure. We argue that this harness is itself a learnable control layer. We formalize harness ope…

  145. arXiv cs.LG TIER_1 English(EN) · Kevin Wang, Kevin Yang, Arjun Prakash, Amy Greenwald ·

    Towards Learning Representations of Policies in Two-Player Zero-Sum Imperfect-Information Games

    arXiv:2607.01498v1 Announce Type: new Abstract: We investigate the problem of learning useful policy representations (embeddings) in two-player zero-sum imperfect-information games. We make three contributions: First, we introduce methods of creating datasets of policies for a gi…

  146. arXiv cs.LG TIER_1 English(EN) · Jen-Yen Chang, Takayuki Osa, Tatsuya Harada ·

    Learning the Supports for Categorical Critic in Reinforcement Learning

    arXiv:2607.01880v1 Announce Type: new Abstract: Value functions are an essential component in actor-critic based deep reinforcement learning (RL). Conventionally, these functions are trained as a regression task by minimising the mean squared error (MSE) relative to bootstrapped …

  147. arXiv cs.AI TIER_1 English(EN) · Yilie Huang, Wenpin Tang, Xun Yu Zhou ·

    ART for Diffusion Sampling: Continuous-Time Control and Actor-Critic Learning

    arXiv:2607.02137v1 Announce Type: cross Abstract: We study timestep allocation for score-based diffusion sampling, where a learned reverse-time dynamics is discretized on a finite grid. Uniform and hand-crafted schedules are standard choices, but they rely on fixed prescriptions …

  148. arXiv cs.AI TIER_1 English(EN) · Yuriy Maksyuta, George Bredis, Ruslan Rakhimov, Daniil Gavrilov ·

    Rank-Then-Act: Reward-Free Control from Frame-Order Progress

    arXiv:2607.01897v1 Announce Type: cross Abstract: We introduce Rank-Then-Act (RTA), a framework for learning control policies from expert video demonstrations without environment rewards. RTA trains a Vision-Language Model (VLM) offline as a progress-based ordinal scorer, using a…

  149. arXiv cs.AI TIER_1 English(EN) · Shenghui Zhang, YuXuan Gao, Songwei Zhao, Jifeng Hu, Zijing Zhang, Hechang Chen ·

    Lightweight Safe Reinforcement Learning for End-to-End UAV Navigation

    arXiv:2607.01794v1 Announce Type: cross Abstract: With the rapid development of autonomous aerial systems, Unmanned Aerial Vehicles (UAVs) are increasingly deployed in applications such as inspection, environmental monitoring, and rescue, creating growing demand for reliable auto…

  150. arXiv cs.AI TIER_1 English(EN) · Juliette Decugis, Sean O'Brien, Francis Bach, Gabriel Synnaeve, Taco Cohen ·

    Don't Let Gains FADE: Breaking Down Policy Gradient Weights in RL

    arXiv:2607.01490v1 Announce Type: cross Abstract: Reinforcement learning post-training dramatically improves LLM reasoning, but suffers from training instability and diversity collapse. Advantage functions offer an appealing fix: they reshape the training objective, reweight whic…

  151. arXiv cs.LG TIER_1 English(EN) · Ren\'e Carmona, Mathieu Lauri\`ere ·

    Mean Field Reinforcement Learning

    arXiv:2607.01525v1 Announce Type: cross Abstract: This monograph provides an introduction to mean field reinforcement learning through the lens of Markov decision processes arising from large-population stochastic control with mean field interactions and common noise. Starting fr…

  152. arXiv cs.AI TIER_1 English(EN) · Anish Diwan, Davide Tateo, Christopher E. Mower, Haitham Bou-Ammar, Jan Peters, Oleg Arenz ·

    Trust Region Inverse Reinforcement Learning: Explicit Dual Ascent using Local Policy Updates

    arXiv:2605.11020v2 Announce Type: replace-cross Abstract: Inverse reinforcement learning (IRL) is typically formulated as maximizing entropy subject to matching the distribution of expert trajectories. Classical (dual-ascent) IRL guarantees monotonic performance improvement but r…

  153. arXiv cs.LG TIER_1 English(EN) · Daniel Thi Graviet, Lovre Pesut, Ivan Dagelic, Vedran Jukic, Ivan Burazin ·

    The Rollout Infrastructure Tax in Coding-Agent Reinforcement Learning

    arXiv:2607.01415v1 Announce Type: new Abstract: Coding-agent reinforcement learning treats execution infrastructure as a background implementation detail, despite relying on large numbers of interactive software rollouts. This is a missed opportunity: measuring infrastructure ove…

  154. arXiv cs.AI TIER_1 English(EN) · Xun Yu Zhou ·

    ART for Diffusion Sampling: Continuous-Time Control and Actor-Critic Learning

    We study timestep allocation for score-based diffusion sampling, where a learned reverse-time dynamics is discretized on a finite grid. Uniform and hand-crafted schedules are standard choices, but they rely on fixed prescriptions and can therefore be suboptimal. To address this l…

  155. Hugging Face Daily Papers TIER_1 English(EN) ·

    ART for Diffusion Sampling: Continuous-Time Control and Actor-Critic Learning

    We study timestep allocation for score-based diffusion sampling, where a learned reverse-time dynamics is discretized on a finite grid. Uniform and hand-crafted schedules are standard choices, but they rely on fixed prescriptions and can therefore be suboptimal. To address this l…

  156. arXiv cs.LG TIER_1 English(EN) · Jinwen Wang, Youfang Lin, Xiaobo Hu, Shuo Wang, Kai Lv ·

    Local Motion Matters: A Deconstruct-Recompose Paradigm for Reinforcement Learning Pre-training from Videos

    arXiv:2607.00808v1 Announce Type: new Abstract: Pre-training on large-scale videos to improve reinforcement learning efficiency is promising yet remains challenging. Existing methods typically treat the agent as an indivisible entity, modeling motion patterns globally. Such globa…

  157. arXiv cs.AI TIER_1 English(EN) · Zishang Jiang, Jinyi Han, Tingyun Li, Xinyi Wang, Sihang Jiang, Jiaqing Liang, Zhaoqian Dai, Shuguang Ma, Fei Yu, Yanghua Xiao ·

    Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs

    arXiv:2510.04140v2 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become a widely adopted technique for enhancing the reasoning ability of Large Language Models (LLMs). However, the effectiveness of RLVR strongly depends on the capabili…

  158. arXiv cs.AI TIER_1 English(EN) · Zhiqi Li, Wen Zhang, Bo Zhu ·

    Flow-Map GRPO: Reinforcement Learning for Few-Step Flow-Map Generators via Anchored Stochastic Composition

    arXiv:2607.00535v1 Announce Type: cross Abstract: Few-step flow-map generators, such as consistency models and MeanFlow, accelerate sampling by directly learning long-range transport maps between noise and data. However, these models are typically deterministic, which makes them …

  159. arXiv cs.AI TIER_1 English(EN) · Konstantin Garbers ·

    Gauging, Measuring, and Controlling Critic Complexity in Actor-Critic Reinforcement Learning

    arXiv:2607.00452v1 Announce Type: cross Abstract: Actor-critic methods depend on learned critics, but critic quality is often evaluated only indirectly through return, temporal-difference error, or value loss. Critic complexity is introduced as an additional diagnostic and interv…

  160. arXiv cs.AI TIER_1 English(EN) · Jongchan Park, Seungjun Oh, Seungho Baek, Yusung Kim ·

    Learning Generalizable Skill Policy with Data-Efficient Unsupervised RL

    arXiv:2607.00392v1 Announce Type: cross Abstract: Unsupervised Reinforcement Learning (URL) aims to pre-train scalable, skill-conditioned policies without extrinsic rewards, serving as a foundation for downstream control tasks. Despite recent progress, we argue that current off-p…

  161. arXiv cs.AI TIER_1 English(EN) · Egor Cherepanov, Daniil Zelezetsky, Alexey K. Kovalev, Aleksandr I. Panov ·

    KAGE-Bench: Fast Known-Axis Visual Generalization Evaluation for Reinforcement Learning

    arXiv:2601.14232v2 Announce Type: replace-cross Abstract: Pixel-based reinforcement learning agents often fail under purely visual distribution shift even when latent dynamics and rewards are unchanged, but existing benchmarks entangle multiple sources of shift and hinder systema…

  162. arXiv cs.LG TIER_1 English(EN) · Gaia Molinaro, Anne G. E. Collins ·

    Reward function compression facilitates goal-dependent reinforcement learning

    arXiv:2509.06810v3 Announce Type: replace-cross Abstract: Humans can uniquely assign value to novel, abstract outcomes to support reinforcement learning. However, this flexibility is cognitively costly and reduces learning efficiency. We propose that goal-dependent learning initi…

  163. arXiv cs.CL TIER_1 English(EN) · Ziyou Hu, Zhengliang Shi, Minghang Zhu, Haitao Li, Teng Sun, Pengjie Ren, Suzan Verberne, Zhaochun Ren ·

    OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning

    arXiv:2510.24636v3 Announce Type: replace Abstract: Reward models (RMs) have become essential for aligning large language models (LLMs), serving as scalable proxies for human evaluation in both training and inference. However, existing RMs struggle on knowledge-intensive and long…

  164. arXiv cs.LG TIER_1 English(EN) · Jinwen Wang, Youfang Lin, Xiaobo Hu, Qian Xu, Shuo Wang, Zhuo Chen, Kai Lv ·

    Task-Relevant Representation Decoupling for Visual Reinforcement Learning Generalization

    arXiv:2607.00796v1 Announce Type: new Abstract: Visual Reinforcement Learning (VRL) has achieved considerable success in solving control tasks. However, generalizing learned policies to new environments remains a major challenge, as agents often overfit to task-irrelevant feature…

  165. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Mathieu Laurière ·

    Mean Field Reinforcement Learning

    This monograph provides an introduction to mean field reinforcement learning through the lens of Markov decision processes arising from large-population stochastic control with mean field interactions and common noise. Starting from the connection between multi-agent reinforcemen…

  166. arXiv cs.LG TIER_1 English(EN) · Kai Lv ·

    Local Motion Matters: A Deconstruct-Recompose Paradigm for Reinforcement Learning Pre-training from Videos

    Pre-training on large-scale videos to improve reinforcement learning efficiency is promising yet remains challenging. Existing methods typically treat the agent as an indivisible entity, modeling motion patterns globally. Such global modeling is tightly coupled with the morpholog…

  167. arXiv cs.LG TIER_1 English(EN) · Kai Lv ·

    Task-Relevant Representation Decoupling for Visual Reinforcement Learning Generalization

    Visual Reinforcement Learning (VRL) has achieved considerable success in solving control tasks. However, generalizing learned policies to new environments remains a major challenge, as agents often overfit to task-irrelevant features in the training environment. To solve this pro…

  168. arXiv cs.AI TIER_1 English(EN) · Bo Zhu ·

    Flow-Map GRPO: Reinforcement Learning for Few-Step Flow-Map Generators via Anchored Stochastic Composition

    Few-step flow-map generators, such as consistency models and MeanFlow, accelerate sampling by directly learning long-range transport maps between noise and data. However, these models are typically deterministic, which makes them difficult to optimize with reinforcement learning …

  169. arXiv cs.AI TIER_1 English(EN) · Konstantin Garbers ·

    Gauging, Measuring, and Controlling Critic Complexity in Actor-Critic Reinforcement Learning

    Actor-critic methods depend on learned critics, but critic quality is often evaluated only indirectly through return, temporal-difference error, or value loss. Critic complexity is introduced as an additional diagnostic and intervention dimension for actor-critic reinforcement le…

  170. arXiv cs.AI TIER_1 English(EN) · Debmalya Mandal, Andi Nika, Parameswaran Kamalaruban, Adish Singla, Goran Radanovi\'c ·

    Corruption Robust Offline Reinforcement Learning with Human Feedback

    arXiv:2402.06734v2 Announce Type: replace-cross Abstract: We study data corruption robustness for reinforcement learning with human feedback (RLHF) in an offline setting. Given an offline dataset of pairs of trajectories along with feedback about human preferences, an $\varepsilo…

  171. arXiv cs.AI TIER_1 English(EN) · Yuanda Xu, Zhengze Zhou, Hejian Sang, Xiaomin Li, Jiaxin Zhang, Xinchen Du, Zhipeng Wang, Alborz Geramifard ·

    TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning

    arXiv:2606.32017v1 Announce Type: cross Abstract: Agentic reinforcement learning requires assigning credit to environment-facing actions such as searches, clicks, edits, navigation commands, and object interactions. Standard GRPO uses the final verifier outcome as a uniform advan…

  172. arXiv cs.AI TIER_1 English(EN) · Arshia Rafieioskouei, Tzu-Han Hsu, Matthew Lucas, Borzoo Bonakdarpour ·

    HyPOLE: Hyperproperty-Guided Multi-Agent Reinforcement Learning under Partial Observation

    arXiv:2606.30966v1 Announce Type: new Abstract: Formal specification is a powerful tool to guide the learning process and provides significant advantages over reward shaping: (1) mathematical rigor; (2) expressiveness to specify objectives and constraints, and (3) the ability to …

  173. arXiv cs.LG TIER_1 English(EN) · Deepak Akhare, Luning Sun, Xin-Yang Liu, Xiantao Fan, Timo Bremer, Ben Zhu, Jian-Xun Wang ·

    Offline Reinforcement Learning for Fluid Controls: Data-based Multi-observational Policy Extraction

    arXiv:2606.31025v1 Announce Type: new Abstract: Active flow control is a fundamental application in engineering. Recent advances in deep reinforcement learning have made progress in this field. However, the classical online RL approaches require extensive real-time interactions w…

  174. arXiv cs.LG TIER_1 English(EN) · Mingyi Li, Taira Tsuchiya, Kenji Yamanishi ·

    Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions

    arXiv:2606.31769v1 Announce Type: new Abstract: We study policy optimization for online episodic tabular Markov decision processes with unknown transition kernels, aiming for best-of-both-worlds guarantees together with data-dependent regret bounds. Recent work (Dann et al., 2023…

  175. arXiv cs.AI TIER_1 English(EN) · Yang Yang, Bingjie Chen, Zihan Wang, Yizhe Li, Guoping Pan, Yi Cheng, Houde Liu ·

    Stage-Transition Dense Reward Modeling for Reinforcement Learning

    arXiv:2606.31377v1 Announce Type: cross Abstract: Reinforcement learning for long-horizon robotic manipulation is often limited by sparse and delayed rewards, while manually designing dense shaping signals is costly and brittle to changes in environments and object configurations…

  176. arXiv cs.LG TIER_1 English(EN) · Ethan Hirschowitz, Fabio Ramos ·

    Warp RL: Reshaping Base Policy Distributions for Dynamics Adaptation

    arXiv:2606.31043v1 Announce Type: new Abstract: Residual reinforcement learning adapts a pretrained robot policy by learning an additive correction to its actions. While effective when adaptation amounts to shifting the base policy's action distribution, additive corrections cann…

  177. arXiv cs.LG TIER_1 English(EN) · Tommaso Marzi, Cesare Alippi, Andrea Cini ·

    Hierarchical Message-Passing Policies for Multi-Agent Reinforcement Learning

    arXiv:2507.23604v2 Announce Type: replace Abstract: Decentralized Multi-Agent Reinforcement Learning (MARL) methods allow for learning scalable multi-agent policies, but suffer from partial observability and induced non-stationarity. These challenges can be addressed by introduci…

  178. arXiv cs.LG TIER_1 English(EN) · Yusung Kim ·

    Learning Generalizable Skill Policy with Data-Efficient Unsupervised RL

    Unsupervised Reinforcement Learning (URL) aims to pre-train scalable, skill-conditioned policies without extrinsic rewards, serving as a foundation for downstream control tasks. Despite recent progress, we argue that current off-policy URL methods are limited by two critical, ove…

  179. arXiv cs.AI TIER_1 English(EN) · Alborz Geramifard ·

    TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning

    Agentic reinforcement learning requires assigning credit to environment-facing actions such as searches, clicks, edits, navigation commands, and object interactions. Standard GRPO uses the final verifier outcome as a uniform advantage over all action tokens. This outcome signal i…

  180. arXiv cs.LG TIER_1 English(EN) · Kenji Yamanishi ·

    Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions

    We study policy optimization for online episodic tabular Markov decision processes with unknown transition kernels, aiming for best-of-both-worlds guarantees together with data-dependent regret bounds. Recent work (Dann et al., 2023; Li et al., 2026) has shown that policy optimiz…

  181. arXiv cs.AI TIER_1 English(EN) · Qijun Li, Zheng Fu, Qi Song, Yifei He, Weitao Zhou, Kun Jiang, Diange Yang ·

    Dual-Flow Reinforcement Learning with State-Aware Exploration

    arXiv:2606.29820v1 Announce Type: cross Abstract: In complex continuous-control reinforcement learning tasks, multimodal optimal actions often coincide with uncertain, multimodal return distributions, making reliable value estimation and multimodal exploration challenging. Existi…

  182. arXiv cs.AI TIER_1 English(EN) · Runze Zhao, Dongruo Zhou, Sumit Kumar Jha, Nathaniel D. Bastian, Ankit Shah ·

    HiComm: Hierarchical Communication for Multi-agent Reinforcement Learning

    arXiv:2606.29126v1 Announce Type: new Abstract: Cooperative multi-agent reinforcement learning (MARL) often relies on communication to mitigate partial observability, yet most existing protocols treat messages as flat dense vectors detached from the structure of the observations …

  183. arXiv cs.AI TIER_1 English(EN) · Daiki E. Matsunaga, Junho Na, Tri Wahyu Guntara, Scott Sanner, Pascal Poupart, Jongmin Lee, Kee-Eung Kim ·

    ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning

    arXiv:2606.30072v1 Announce Type: new Abstract: Cooperative tasks in Multi-Agent Reinforcement Learning (MARL) require agents to collectively maximize a shared return. Under the Centralized Training with Decentralized Execution (CTDE) paradigm, policy gradients have remained diff…

  184. arXiv cs.AI TIER_1 English(EN) · Adithya Mohan, Daniel Kriegl, Torsten Sch\"on ·

    RoAd-RL: A Unified Library and Benchmark for Robust Adversarial Reinforcement Learning

    arXiv:2606.29867v1 Announce Type: cross Abstract: Deep Reinforcement Learning (DRL) has achieved significant success in robotics and autonomous systems, yet remains vulnerable to adversarial perturbations that can severely degrade performance. Research in adversarial reinforcemen…

  185. arXiv cs.AI TIER_1 English(EN) · Zibin Meng, Kani Chen ·

    CRAFT: Counterfactual Credit Assignment from Free Sibling Rollouts for Self-Distilled Agentic Reinforcement Learning

    arXiv:2606.29476v1 Announce Type: cross Abstract: Self-distilled agentic reinforcement learning augments trajectory-level reward with a token-level distillation loss, using as its teacher the same policy conditioned on privileged context. The prevailing recipe gates this loss by …

  186. arXiv cs.LG TIER_1 English(EN) · Zheming Fu, Ruizhe He, Wei Shang, Xiaoxiao Ma, Lei Wang, Chang Liu, Siming Fu ·

    FlowAWR: Online Adaptive Flow Reinforcement via Advantage-Weighted Rectification

    arXiv:2606.30376v1 Announce Type: new Abstract: Aligning generative flow models on continuous spaces via online reinforcement learning is constrained by intractable trajectory likelihoods. Existing density-approximated policy gradient methods rely on stochastic SDE samplers to co…

  187. arXiv cs.LG TIER_1 English(EN) · Jing Liang, Hongyao Tang, Yi Ma, Yancheng He, Weixun Wang, Xiaoyang Li, Ju Huang, Wenbo Su, Jinyi Liu, Yan Zheng, Jianye Hao, Bo Zheng ·

    The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning

    arXiv:2606.29526v1 Announce Type: new Abstract: Reinforcement learning (RL) has gained growing attention in large language model (LLM) post-training, yet RL training remains fragile and can suffer from instability or collapse. One vital cause is training-inference mismatch: LLM a…

  188. arXiv cs.LG TIER_1 English(EN) · Jesse Ponnock, Lucas Ho ·

    Reinforcement Learning in Super Mario Bros: Curriculum, Pedagogy, and Optimal Level Design in World 1-1

    arXiv:2606.29511v1 Announce Type: new Abstract: World 1-1 of Super Mario Bros is widely celebrated as a masterclass in game design: its progressive structure is credited with teaching players core mechanics through the level itself. We ask whether that structure is empirically me…

  189. arXiv cs.AI TIER_1 English(EN) · Hao Wang, Jiuzhou Lei, Dayou Li, Bangya Liu, Minghui Zheng, Manling Li, Ruohan Zhang, Zhiwen Fan ·

    Behavior Uncloning: Distilling Mode Redirection into Policy Weights without Inference-Time Steering

    arXiv:2606.29201v1 Announce Type: cross Abstract: Behavior-cloned policies often learn multiple behavior modes from demonstration datasets, including modes that are unsafe or otherwise undesired at deployment. For example, a policy trained on diverse handover demonstrations may l…

  190. arXiv cs.LG TIER_1 English(EN) · Amritansh Mishra, Supriyo Chakraborty, Berkcan Kapusuzoglu ·

    On the Policy Gradient Foundations of Group Relative Policy Optimization: Credit Assignment, Gradient Sparsity, and Rank Collapse

    arXiv:2606.29238v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) eliminates the learned critic in PPO by using the mean reward of grouped rollouts as a baseline. We provide a rigorous derivation of GRPO from first principles of the policy gradient theorem…

  191. arXiv cs.LG TIER_1 English(EN) · Congde Hu, Danping Li, Lin Xu, Wenying Xu ·

    Entropy-Regularized Reinforcement Learning for Linear-Quadratic Stackelberg Differential Games in Regime-Switching Diffusion Models

    arXiv:2606.28671v1 Announce Type: new Abstract: Stackelberg differential games (SDGs) provide a powerful framework for hierarchical decision-making in stochastic and continuous-time environments, yet their solution remains computationally challenging due to the complexity of trad…

  192. arXiv cs.LG TIER_1 English(EN) · Congde Hu, Zhuo Jin, Danping Li, Lin Xu ·

    Entropy Regularized Reinforcement Learning for Zero-Sum Stochastic Differential Games in a Regime-Switching Jump-Diffusion Process

    arXiv:2606.28669v1 Announce Type: new Abstract: To address parameter misspecification and sudden structural environmental changes in conventional stochastic differential game (SDG) frameworks, this paper introduces a distributional control approach that characterizes optimal stra…

  193. arXiv cs.AI TIER_1 English(EN) · Rajib Mostakim, Reza T. Batley, Sourav Saha ·

    Agile Reinforcement Learning through Separable Neural Architecture and Applications

    arXiv:2601.23225v2 Announce Type: replace-cross Abstract: Deep reinforcement learning (RL) is increasingly deployed in resource-constrained environments, yet go-to function approximators - multilayer perceptrons (MLPs) - are often parameter-inefficient due to an imperfect inducti…

  194. arXiv cs.AI TIER_1 English(EN) · Mathieu Petitbois, R\'emy Portelas, Sylvain Lamprier ·

    Offline Reinforcement Learning of High-Quality Behaviors Under Robust Style Alignment

    arXiv:2601.22823v2 Announce Type: replace-cross Abstract: We study offline reinforcement learning of style-conditioned policies using explicit style supervision via subtrajectory labeling functions. In this setting, aligning style with high task performance is particularly challe…

  195. arXiv cs.AI TIER_1 English(EN) · Debmalya Mandal, Paulius Sasnauskas, Goran Radanovic ·

    Distributionally Robust Reinforcement Learning with Human Feedback

    arXiv:2503.00539v2 Announce Type: replace-cross Abstract: Reinforcement learning from human feedback (RLHF) has evolved to be one of the main methods for fine-tuning large language models (LLMs). However, existing RLHF methods are non-robust, and their performance deteriorates if…

  196. arXiv cs.AI TIER_1 English(EN) · Charles Westphal, Stephen Hailes, Mirco Musolesi ·

    TERC: A Transfer Entropy Redundancy Criterion for State Variable Selection in Reinforcement Learning

    arXiv:2401.11512v2 Announce Type: replace-cross Abstract: Identifying the most suitable variables to represent the state is a fundamental challenge in Reinforcement Learning (RL). These variables must efficiently capture the information necessary for making optimal decisions. In …

  197. Hugging Face Daily Papers TIER_1 English(EN) ·

    TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning

    TRIAGE introduces a role-typed credit assignment framework that enhances agentic reinforcement learning by providing more nuanced credit assignment than standard GRPO methods.

  198. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Borzoo Bonakdarpour ·

    HyPOLE: Hyperproperty-Guided Multi-Agent Reinforcement Learning under Partial Observation

    Formal specification is a powerful tool to guide the learning process and provides significant advantages over reward shaping: (1) mathematical rigor; (2) expressiveness to specify objectives and constraints, and (3) the ability to define tactics to achieve objectives. However, t…

  199. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Claudio Pacchierotti ·

    Sampling-Based Coordination-Informed Multi-Objective Multi-Robot Reinforcement Learning

    Multi-robot systems must simultaneously optimize competing objectives while maintaining coordinated behavior. Existing multi-agent reinforcement learning approaches often rely on fixed or centralized coordination, which limits adaptability and violates distributed constraints. Th…

  200. arXiv cs.LG TIER_1 English(EN) · Siming Fu ·

    FlowAWR: Online Adaptive Flow Reinforcement via Advantage-Weighted Rectification

    Aligning generative flow models on continuous spaces via online reinforcement learning is constrained by intractable trajectory likelihoods. Existing density-approximated policy gradient methods rely on stochastic SDE samplers to construct tractable transition kernels, which intr…

  201. arXiv cs.AI TIER_1 English(EN) · Difan Jiao, Raghav Singhal, Robert West, Ashton Anderson ·

    Tandem Reinforcement Learning with Verifiable Rewards

    arXiv:2606.28166v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has significantly improved the reasoning capability of large language models, reaching expert or even superhuman performance in domains such as competition math. However, whether…

  202. arXiv cs.AI TIER_1 English(EN) · Qitai Tan, Zefang Zong, Yang Li, Peng Chen ·

    ATOD: Annealed Turn-aware On-policy Distillation for Multi-turn Autonomous Agents

    arXiv:2606.27814v1 Announce Type: new Abstract: Training small language-model agents for long-horizon interactive tasks requires both fast imitation and reward-driven improvement. On-policy distillation (OPD) provides dense teacher guidance and typically improves rapidly in the e…

  203. arXiv cs.LG TIER_1 English(EN) · Jiexin Wang, Eiji Uchibe ·

    Regularized Reward-Punishment Reinforcement Learning

    arXiv:2606.28152v1 Announce Type: new Abstract: We propose KL-Coupled Policy Regularization (KCPR), a policy coordination framework for Reward-Punishment Reinforcement Learning (RPRL). Based on KCPR, we derive KL-Coupled Soft Optimality (KCSO) and develop its deep realization, kl…

  204. arXiv cs.LG TIER_1 English(EN) · Tianlin Pan, Lianyu Pang, Cheng Da, Huan Yang, Changqian Yu, Kun Gai, Wenhan Luo ·

    NormGuard: Reward-Preserving Norm Constraints in Flow-Matching Reinforcement Learning

    arXiv:2606.27771v1 Announce Type: new Abstract: Reinforcement learning (RL) post-training improves the reward alignment of flow-based generators, but often degrades perceptual quality in ways that are not captured by the reward proxy. We identify a simple structural signature of …

  205. arXiv cs.AI TIER_1 English(EN) · Austin A. Nguyen, Michael P. Wellman ·

    Conservative Equilibrium Discovery in Offline Game-Theoretic Multiagent Reinforcement Learning

    arXiv:2603.00374v2 Announce Type: replace Abstract: Offline learning of strategies takes data efficiency to its extreme by restricting algorithms to a fixed dataset of state-action trajectories. We consider the problem in a mixed-motive multiagent setting, where the goal is to so…

  206. Hugging Face Daily Papers TIER_1 English(EN) ·

    The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning

    Training-inference mismatch in reinforcement learning for large language models leads to instability, which is addressed through a new policy optimization objective and framework that ensures consistent policy improvements between training and inference phases.

  207. arXiv cs.AI TIER_1 English(EN) · Ashton Anderson ·

    Tandem Reinforcement Learning with Verifiable Rewards

    Reinforcement learning with verifiable rewards (RLVR) has significantly improved the reasoning capability of large language models, reaching expert or even superhuman performance in domains such as competition math. However, whether weaker agents and humans can actually harness t…

  208. arXiv cs.LG TIER_1 English(EN) · Eiji Uchibe ·

    Regularized Reward-Punishment Reinforcement Learning

    We propose KL-Coupled Policy Regularization (KCPR), a policy coordination framework for Reward-Punishment Reinforcement Learning (RPRL). Based on KCPR, we derive KL-Coupled Soft Optimality (KCSO) and develop its deep realization, klDMP. Unlike existing RPRL approaches that optimi…

  209. arXiv cs.AI TIER_1 English(EN) · Peng Chen ·

    ATOD: Annealed Turn-aware On-policy Distillation for Multi-turn Autonomous Agents

    Training small language-model agents for long-horizon interactive tasks requires both fast imitation and reward-driven improvement. On-policy distillation (OPD) provides dense teacher guidance and typically improves rapidly in the early stage, but its gains saturate once the stud…

  210. arXiv cs.LG TIER_1 English(EN) · Wenhan Luo ·

    NormGuard: Reward-Preserving Norm Constraints in Flow-Matching Reinforcement Learning

    Reinforcement learning (RL) post-training improves the reward alignment of flow-based generators, but often degrades perceptual quality in ways that are not captured by the reward proxy. We identify a simple structural signature of this drift: across three post-training methods (…

  211. arXiv cs.LG TIER_1 English(EN) · Wenjie Huang, Yang Li, Jingjia Teng, Mingwei Jin, Kai Song, Yougang Bian, Yongfu Li, Qisong Yang, Helai Huang ·

    Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy

    arXiv:2606.26527v1 Announce Type: new Abstract: Transfer learning improves policy learning efficiency by reusing knowledge from source tasks, providing a feasible paradigm for safe and efficient autonomous highway lane changing decision-making. Existing methods frequently encount…

  212. arXiv cs.LG TIER_1 English(EN) · Behnam Gheshlaghi, Bahador Rashidi, Shahin Atakishiyev ·

    Mesh-RL: Coupled subgrid reinforcement learning

    arXiv:2606.26333v1 Announce Type: new Abstract: Reinforcement learning in large or sparse-reward environments suffers from slow temporal-difference reward propagation, as value information spreads only locally across the state space. We propose Mesh-RL, a spatial domain-decomposi…

  213. arXiv cs.AI TIER_1 English(EN) · Aniruddha Joshi, Niklas Lauffer, Sanjit Seshia ·

    Deterministic Pareto-Optimal Policy Synthesis for Multi-Objective Reinforcement Learning

    arXiv:2606.26397v1 Announce Type: cross Abstract: Real-world decision-making often requires balancing multiple conflicting objectives, a challenge that standard Reinforcement Learning (RL) frequently addresses by aggregating rewards into a single scalar signal. While effective fo…

  214. arXiv cs.AI TIER_1 English(EN) · Ting Zhou, Zhenqing Ling, Yiyang Zhao, Ying Shen, Daoyuan Chen ·

    GEOALIGN: Geometric Rollout Curation for Robust LLM Reinforcement Learning

    arXiv:2606.26917v1 Announce Type: cross Abstract: Online reinforcement learning is widely used to align large language models (LLMs) with reward signals, yet training can be unstable under noisy or misspecified rewards. We identify a failure mode we call directional inconsistency…

  215. arXiv cs.AI TIER_1 English(EN) · Jesper Klicks, Sander Vr\v{z}ina, Vincent Fran\c{c}ois-Lavet ·

    State Representation Matters in Deep Reinforcement Learning: Application to Energy Trading

    arXiv:2606.27032v1 Announce Type: cross Abstract: Energy trading decisions depend not only on current market prices, but also on expected future market conditions, and operational constraints. This makes the state representation given to a reinforcement learning agent an importan…

  216. arXiv cs.CL TIER_1 English(EN) · Shuo Yang, Jinyang Wu, Zhengxi Lu, Yuhao Shen, Fan Zhang, Lang Feng, Shuai Zhang, Haoran Luo, Zheng Lian, Zhengqi Wen, Jianhua Tao ·

    OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning

    arXiv:2606.26790v1 Announce Type: new Abstract: Outcome-based reinforcement learning provides a stable optimization backbone for language agents, but its sparse trajectory-level rewards provide little guidance on which intermediate decisions should be reinforced or suppressed. On…

  217. arXiv cs.LG TIER_1 English(EN) · Rongjian Chen, Jianmin Hu, Kejiang Ye, Minxian Xu ·

    RolloutPipe: Overlapping Pipelined Rollout and Training in Disaggregated On-Policy LLM Reinforcement Learning

    arXiv:2606.26997v1 Announce Type: cross Abstract: Large language model (LLM) post-training for reasoning increasingly relies on reinforcement learning with verifiable rewards (RLVR), where models learn from ground-truth feedback on mathematical, logical, and scientific tasks. To …

  218. arXiv cs.LG TIER_1 English(EN) · Erhan Bayraktar, Martin Hernandez, Qinxin Yan, Yuhua Zhu ·

    Mean-Field PhiBE: Continuous-Time Mean-Field Reinforcement Learning from Discrete-Time Data

    arXiv:2606.26498v1 Announce Type: cross Abstract: This paper addresses model-free continuous-time mean-field control in a setting where the population dynamics evolve continuously according to an unknown McKean-Vlasov stochastic differential equation, while only discrete-time tra…

  219. arXiv cs.LG TIER_1 English(EN) · Yingyu Lin, Qiyue Gao, Nikki Lijing Kuang, Xunpeng Huang, Kun Zhou, Tongtong Liang, Zhewei Yao, Yi-An Ma, Yuxiong He ·

    Reinforcement Learning without Ground-Truth Solutions can Improve LLMs

    arXiv:2606.27369v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) for training LLMs typically rely on ground-truth answers to assign rewards, limiting their applicability to tasks where the ground-truth solution is unknown. We introduce a \text…

  220. arXiv cs.AI TIER_1 English(EN) · Shicheng Ye, Chao Yu ·

    Joint Learning of Experiential Rules and Policies for Large Language Model Agents

    arXiv:2606.27136v1 Announce Type: new Abstract: For LLM agents in multi-step interactive environments, a key challenge is to make effective use of accumulated interaction experience. Existing work has typically separated two uses of such experience: keeping it outside the model a…

  221. arXiv cs.AI TIER_1 English(EN) · Boyun Zhang, Chao Wang, Kai Wu ·

    EVOM: Agentic Meta-Evolution of Actor-Critic Architectures for Reinforcement Learning

    arXiv:2606.26327v1 Announce Type: cross Abstract: In actor-critic reinforcement learning, network architectures are typically manually designed. Automating this design is challenging because each candidate must be trained before evaluation, and the design space is open-ended. To …

  222. Hugging Face Daily Papers TIER_1 English(EN) ·

    NormGuard: Reward-Preserving Norm Constraints in Flow-Matching Reinforcement Learning

    Reinforcement learning post-training degrades perceptual quality in flow-based generators through velocity norm inflation, which requires training-time intervention rather than inference-time corrections to maintain both reward alignment and image quality.

  223. arXiv cs.LG TIER_1 English(EN) · Yuxiong He ·

    Reinforcement Learning without Ground-Truth Solutions can Improve LLMs

    Reinforcement learning with verifiable rewards (RLVR) for training LLMs typically rely on ground-truth answers to assign rewards, limiting their applicability to tasks where the ground-truth solution is unknown. We introduce a \textbf{R}anking-\textbf{i}nduced \textbf{VER}ifiable…

  224. arXiv cs.AI TIER_1 English(EN) · Chao Yu ·

    Joint Learning of Experiential Rules and Policies for Large Language Model Agents

    For LLM agents in multi-step interactive environments, a key challenge is to make effective use of accumulated interaction experience. Existing work has typically separated two uses of such experience: keeping it outside the model as natural-language rules for later prompting, or…

  225. arXiv cs.LG TIER_1 English(EN) · Vincent François-Lavet ·

    State Representation Matters in Deep Reinforcement Learning: Application to Energy Trading

    Energy trading decisions depend not only on current market prices, but also on expected future market conditions, and operational constraints. This makes the state representation given to a reinforcement learning agent an important design choice. We study this in HydroDam, a pump…

  226. arXiv cs.LG TIER_1 English(EN) · Minxian Xu ·

    RolloutPipe: Overlapping Pipelined Rollout and Training in Disaggregated On-Policy LLM Reinforcement Learning

    Large language model (LLM) post-training for reasoning increasingly relies on reinforcement learning with verifiable rewards (RLVR), where models learn from ground-truth feedback on mathematical, logical, and scientific tasks. To enable flexible resource allocation and support he…

  227. arXiv cs.LG TIER_1 English(EN) · Daoyuan Chen ·

    GEOALIGN: Geometric Rollout Curation for Robust LLM Reinforcement Learning

    Online reinforcement learning is widely used to align large language models (LLMs) with reward signals, yet training can be unstable under noisy or misspecified rewards. We identify a failure mode we call directional inconsistency: within a batch, a small set of high-reward rollo…

  228. arXiv cs.CL TIER_1 English(EN) · Jianhua Tao ·

    OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning

    Outcome-based reinforcement learning provides a stable optimization backbone for language agents, but its sparse trajectory-level rewards provide little guidance on which intermediate decisions should be reinforced or suppressed. On-policy self-distillation offers dense token-lev…

  229. arXiv cs.CL TIER_1 English(EN) · Renfei Zhang, Manasa Kaniselvan, Rylan Schaeffer abd Niloofar Mireshghallah ·

    Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs

    arXiv:2511.05933v2 Announce Type: replace Abstract: Reinforcement learning (RL) is often credited with improving language model reasoning at the expense of knowledge. We challenge this narrative by showing that reasoning models consistently outperform their instruction-tuned vers…

  230. arXiv cs.LG TIER_1 English(EN) · Yuexin Bian, Jie Feng, Tao Wang, Yijiang Li, Sicun Gao, Yuanyuan Shi ·

    RN-D: Discretized Categorical Actors for On-Policy Reinforcement Learning

    arXiv:2601.23075v2 Announce Type: replace Abstract: On-policy Reinforcement Learning (RL) remains a dominant paradigm for continuous control, yet standard implementations rely on Gaussian actors and relatively shallow MLP policies, often leading to brittle optimization when gradi…

  231. arXiv cs.LG TIER_1 English(EN) · Caleb Ju, Guanghui Lan ·

    Auto-exploration for online reinforcement learning

    arXiv:2512.06244v2 Announce Type: replace Abstract: The exploration-exploitation dilemma in reinforcement learning (RL) is a fundamental challenge to efficient RL algorithms. Existing algorithms for finite state and action discounted RL problems address this by assuming sufficien…

  232. arXiv cs.LG TIER_1 English(EN) · Thiago Thomas, Gabriel de Oliveira Ramos, Felipe Meneguzzi ·

    Multi-Agent Goal Recognition with Team- and Goal-Conditioned Reinforcement Learning and Factorized Branch-and-Bound

    arXiv:2606.25978v1 Announce Type: cross Abstract: Multi-agent goal recognition asks an observer to jointly infer which agents act together and what each team is trying to achieve, so the hypothesis space grows combinatorially with the number of team partitions and goals per team.…

  233. arXiv cs.LG TIER_1 English(EN) · Peng Xu, Sijia Chen, Junzhuo Li, Xuming Hu ·

    Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents

    arXiv:2606.25852v1 Announce Type: new Abstract: Group-based reinforcement learning effectively post-trains LLM agents for long-horizon, sparse-reward tasks by deriving step-level credit from trajectory outcomes. However, this ties a step's credit to its rollout's final outcome: s…

  234. arXiv cs.LG TIER_1 English(EN) · Samuel Valland Lyngset, Tor Viljen Raanaas, Gard Sveipe, Eirik M{\o}ller Nilsen, Jim Torresen, Kai Olav Ellefsen, Tobias L{\o}mo ·

    Memory-Efficient Policy Libraries with Low-Rank Adaptation in Reinforcement Learning

    arXiv:2606.25700v1 Announce Type: new Abstract: When fine-tuning Large Language Models (LLMs), there has been success in minimizing both memory usage and computation with Parameter-Efficient Fine-Tuning (PEFT), like Low Rank Adaptation (LoRA). In this article, we have explored wh…

  235. arXiv cs.LG TIER_1 English(EN) · Guozheng Ma, Lu Li, Zilin Wang, Pierre-Luc Bacon, Dacheng Tao ·

    Beyond One-Size-Fits-All: Diagnosis-Driven Online Reinforcement Learning with Offline Priors

    arXiv:2606.25527v1 Announce Type: new Abstract: Online reinforcement learning (RL) agents increasingly depend on knowledge acquired offline to achieve practical efficiency. Originally studied in offline-to-online RL, this paradigm now spans foundation model post-training and embo…

  236. arXiv cs.LG TIER_1 English(EN) · Bang Giang Le, Viet Cuong Ta ·

    Low Variance Trust Region Optimization with Independent Actors and Sequential Updates in Cooperative Multi-agent Reinforcement Learning

    arXiv:2606.25526v1 Announce Type: new Abstract: Cooperative multi-agent reinforcement learning assumes each agent shares the same reward function and can be trained effectively using the Trust Region framework of single-agent. Instead of relying on other agents' actions, the inde…

  237. arXiv cs.LG TIER_1 English(EN) · Yivan Zhang, Ziyan Luo, Manuel Baltieri ·

    Compositional Behavioral Semantics for State Abstraction in Reinforcement Learning

    arXiv:2606.25357v1 Announce Type: new Abstract: State abstraction plays a key role in scaling reinforcement learning to complex but structured systems. In studying such systems, a wide range of behavioral structures have been studied in reinforcement learning, including value fun…

  238. arXiv cs.LG TIER_1 English(EN) · Zhengzhu Liu, Zeming Gao, Haoyuan Qin, Jiawei Hu, Junhao Wu, Miao Zhu, Haipeng Zhang, Chennan Ma, Siqi Shen, Cheng Wang ·

    Stagnant Neuron: Towards Understanding the Plasticity Loss in Multi-Agent Reinforcement Learning Value Factorization Methods

    arXiv:2606.25335v1 Announce Type: new Abstract: Multi-Agent Reinforcement Learning (MARL) value factorization methods can suffer from a loss of plasticity, gradually failing to adapt when transferring to new task instances. We trace this issue to stagnant neurons, units whose gra…

  239. arXiv cs.LG TIER_1 English(EN) · Animesh Animesh, Satheesh K Perepu, Kaushik Dey ·

    GCT-MARL: Graph-Based Contrastive Transfer for Sample-Efficient Cooperative Multi-Agent Reinforcement Learning

    arXiv:2606.25073v1 Announce Type: new Abstract: In cooperative multi-agent reinforcement learning (MARL), from a deployment perspective, it is challenging and expensive to train agents from scratch for each new environment or task. In this work, we propose GCT-MARL, a transfer le…

  240. arXiv cs.LG TIER_1 English(EN) · Ankur Naskar, Swetha Ganesh, Vaneet Aggarwal ·

    Bias-Controlled Primal-Dual Natural Actor-Critic: Optimal Rates for Constrained Multi-Objective Average-Reward RL

    arXiv:2606.25012v1 Announce Type: new Abstract: Many reinforcement learning (RL) problems in the infinite-horizon average-reward setting require optimizing multiple conflicting objectives while satisfying multiple safety constraints. A common approach is concave scalarization, wh…

  241. arXiv cs.LG TIER_1 English(EN) · Wenyang Hu, Junxiang Jia, Zhen Shu, Daniel Dahlmeier, See-Kiong Ng, Bryan Kian Hsiang Low ·

    ExTra: Exploratory Trajectory Optimization for Language Model Reinforcement Learning

    arXiv:2606.24994v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) for language-model reasoning can fail at both extremes of task difficulty: easy prompts often produce all-correct, low-diversity rollout groups with little gradient signal, while…

  242. arXiv cs.LG TIER_1 English(EN) · Thibaut Kulak ·

    Towards Scalable Multi-Task Reinforcement Learning with Large Decision Models

    arXiv:2606.24962v1 Announce Type: new Abstract: Recent progress in large-scale sequence modeling has shown that a single model can learn useful representations across highly diverse data distributions. Inspired by these advances, we investigate whether a unified transformer polic…

  243. arXiv cs.LG TIER_1 English(EN) · Haoyuan Deng, Yihong Zhou, Thomas Morstyn, Yi Wang ·

    Supervised Reinforcement Learning for the Coordination of Distributed Energy Resources

    arXiv:2606.24947v1 Announce Type: new Abstract: The increasing integration of distributed energy resources (DERs) is crucial for power system decarbonization, yet unlocking DERs' flexibility is challenged by their inherent uncertainties and modelling complexity. As traditional op…

  244. arXiv cs.CL TIER_1 English(EN) · Yupu Hao, Zhuoran Jin, Huanxuan Liao, Kang Liu, Jun Zhao ·

    Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It

    arXiv:2606.26027v1 Announce Type: new Abstract: Tool use enables large language models (LLMs) to perform complex tasks, and recent agentic reinforcement learning (RL) methods show promise for enhancing model capabilities. However, RL alone often leads to instability or limited ga…

  245. arXiv cs.CL TIER_1 English(EN) · Wenxuan Jiang, Zining Fan, Zijian Zhang, Xuecheng Wu, Hongming Tan, Haoyang Dai, Xiaoyu Li, Xuezhi Cao, Ninghao Liu ·

    OPERA: Aligning Open-Ended Reasoning via Objective Perplexity-based Reinforcement Learning

    arXiv:2606.25757v1 Announce Type: new Abstract: Reinforcement Learning (RL) has enabled LLMs to excel in objective reasoning tasks such as mathematics and code generation. However, applying RL to open-ended tasks, such as creative writing, remains challenging because LLM-as-a-jud…

  246. arXiv cs.CL TIER_1 English(EN) · Hanyang Wang, Weijieying Ren, Yuxiang Zhang, Ding Cao, Zhizhao Zeng, Ke Zeng, Tianxiang Zhao ·

    BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents

    arXiv:2606.25556v1 Announce Type: new Abstract: Stepwise group-based RL is an attractive way to train long-horizon LLM agents without a learned critic: it reuses multiple sampled rollouts to estimate local advantages. Its weakness is less visible but more fundamental: every group…

  247. arXiv cs.LG TIER_1 English(EN) · Helai Huang ·

    Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy

    Transfer learning improves policy learning efficiency by reusing knowledge from source tasks, providing a feasible paradigm for safe and efficient autonomous highway lane changing decision-making. Existing methods frequently encounter transfer mismatch induced by distribution shi…

  248. Hugging Face Daily Papers TIER_1 English(EN) ·

    Mean-Field PhiBE: Continuous-Time Mean-Field Reinforcement Learning from Discrete-Time Data

    This paper addresses model-free continuous-time mean-field control in a setting where the population dynamics evolve continuously according to an unknown McKean-Vlasov stochastic differential equation, while only discrete-time transition data are available. In the model-based for…

  249. arXiv cs.LG TIER_1 English(EN) · Yuhua Zhu ·

    Mean-Field PhiBE: Continuous-Time Mean-Field Reinforcement Learning from Discrete-Time Data

    This paper addresses model-free continuous-time mean-field control in a setting where the population dynamics evolve continuously according to an unknown McKean-Vlasov stochastic differential equation, while only discrete-time transition data are available. In the model-based for…

  250. Hugging Face Daily Papers TIER_1 English(EN) ·

    OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning

    On-policy skill distillation framework extracts dense hindsight supervision from completed trajectories to improve language agent training efficiency and performance.

  251. arXiv cs.CL TIER_1 English(EN) · Jun Zhao ·

    Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It

    Tool use enables large language models (LLMs) to perform complex tasks, and recent agentic reinforcement learning (RL) methods show promise for enhancing model capabilities. However, RL alone often leads to instability or limited gains in tool-use tasks. In our experiments, some …

  252. Hugging Face Daily Papers TIER_1 English(EN) ·

    Multi-Agent Goal Recognition with Team- and Goal-Conditioned Reinforcement Learning and Factorized Branch-and-Bound

    Multi-agent goal recognition asks an observer to jointly infer which agents act together and what each team is trying to achieve, so the hypothesis space grows combinatorially with the number of team partitions and goals per team. Real applications such as drone surveillance and …

  253. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Felipe Meneguzzi ·

    Multi-Agent Goal Recognition with Team- and Goal-Conditioned Reinforcement Learning and Factorized Branch-and-Bound

    Multi-agent goal recognition asks an observer to jointly infer which agents act together and what each team is trying to achieve, so the hypothesis space grows combinatorially with the number of team partitions and goals per team. Real applications such as drone surveillance and …

  254. Hugging Face Daily Papers TIER_1 English(EN) ·

    Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents

    Group-based reinforcement learning effectively post-trains LLM agents for long-horizon, sparse-reward tasks by deriving step-level credit from trajectory outcomes. However, this ties a step's credit to its rollout's final outcome: semantically near-identical intermediate steps re…

  255. arXiv cs.AI TIER_1 English(EN) · Xuming Hu ·

    Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents

    Group-based reinforcement learning effectively post-trains LLM agents for long-horizon, sparse-reward tasks by deriving step-level credit from trajectory outcomes. However, this ties a step's credit to its rollout's final outcome: semantically near-identical intermediate steps re…

  256. arXiv cs.CL TIER_1 English(EN) · Ninghao Liu ·

    OPERA: Aligning Open-Ended Reasoning via Objective Perplexity-based Reinforcement Learning

    Reinforcement Learning (RL) has enabled LLMs to excel in objective reasoning tasks such as mathematics and code generation. However, applying RL to open-ended tasks, such as creative writing, remains challenging because LLM-as-a-judge reward models often exhibit stylistic biases …

  257. Hugging Face Daily Papers TIER_1 English(EN) ·

    OPERA: Aligning Open-Ended Reasoning via Objective Perplexity-based Reinforcement Learning

    Reinforcement Learning (RL) has enabled LLMs to excel in objective reasoning tasks such as mathematics and code generation. However, applying RL to open-ended tasks, such as creative writing, remains challenging because LLM-as-a-judge reward models often exhibit stylistic biases …

  258. arXiv cs.LG TIER_1 English(EN) · Tobias Lømo ·

    Memory-Efficient Policy Libraries with Low-Rank Adaptation in Reinforcement Learning

    When fine-tuning Large Language Models (LLMs), there has been success in minimizing both memory usage and computation with Parameter-Efficient Fine-Tuning (PEFT), like Low Rank Adaptation (LoRA). In this article, we have explored whether this approach is transferable to the world…

  259. arXiv cs.AI TIER_1 English(EN) · Tianxiang Zhao ·

    BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents

    Stepwise group-based RL is an attractive way to train long-horizon LLM agents without a learned critic: it reuses multiple sampled rollouts to estimate local advantages. Its weakness is less visible but more fundamental: every group-relative estimator assumes that the steps it co…

  260. arXiv cs.LG TIER_1 English(EN) · Dacheng Tao ·

    Beyond One-Size-Fits-All: Diagnosis-Driven Online Reinforcement Learning with Offline Priors

    Online reinforcement learning (RL) agents increasingly depend on knowledge acquired offline to achieve practical efficiency. Originally studied in offline-to-online RL, this paradigm now spans foundation model post-training and embodied intelligence, with prior types expanding fr…

  261. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Viet Cuong Ta ·

    Low Variance Trust Region Optimization with Independent Actors and Sequential Updates in Cooperative Multi-agent Reinforcement Learning

    Cooperative multi-agent reinforcement learning assumes each agent shares the same reward function and can be trained effectively using the Trust Region framework of single-agent. Instead of relying on other agents' actions, the independent actors setting considers each agent to a…

  262. Hugging Face Daily Papers TIER_1 English(EN) ·

    Low Variance Trust Region Optimization with Independent Actors and Sequential Updates in Cooperative Multi-agent Reinforcement Learning

    Cooperative multi-agent reinforcement learning assumes each agent shares the same reward function and can be trained effectively using the Trust Region framework of single-agent. Instead of relying on other agents' actions, the independent actors setting considers each agent to a…

  263. arXiv cs.AI TIER_1 English(EN) · Elias Bareinboim, Junzhe Zhang, Sanghack Lee ·

    An Introduction to Causal Reinforcement Learning

    arXiv:2606.24160v1 Announce Type: new Abstract: Causal inference provides a set of principles and tools that allow one to combine data and knowledge about an environment to reason with questions of counterfactual nature, i.e., what would have happened had reality been different, …

  264. arXiv cs.AI TIER_1 English(EN) · Akshay V. Jagadeesh, Rahul K. Arora, Khaled Saab, Ali Malik, Mikhail Trofimov, Foivos Tsimpourlas, Johannes Heidecke, Karan Singhal ·

    Reinforcement Learning Towards Broadly and Persistently Beneficial Models

    arXiv:2606.24014v1 Announce Type: new Abstract: As AI systems are deployed across increasingly diverse and high-stakes settings, model alignment must generalize beyond the tasks and domains seen during training. This is especially important for reinforcement learning (RL), which …

  265. arXiv cs.AI TIER_1 English(EN) · Tianyuan Shi, Canbin Huang, Bei Li, Xin Chen, Xiaojun Quan, Jingang Wang, Qifan Wang ·

    Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning

    arXiv:2606.24064v1 Announce Type: new Abstract: Distilling reasoning capabilities from strong to weak language models typically involves imitating specific solution trajectories, effectively transferring what to answer rather than how to reason. This trajectory-level imitation en…

  266. arXiv cs.LG TIER_1 English(EN) · Luca Viano, Till Freihaut, Emanuele Nevali, Volkan Cevher, Matthieu Geist, Giorgia Ramponi ·

    Multi-agent imitation learning with function approximation: Linear Markov games and beyond

    arXiv:2602.22810v2 Announce Type: replace Abstract: In this work, we present the first theoretical analysis of multi-agent imitation learning (MAIL) in linear Markov games where both the transition dynamics and each agent's reward function are linear in some given features. We de…

  267. arXiv cs.AI TIER_1 English(EN) · Anurag Akula, Satheesh K. Perepu, Abhishek Sarkar, Kaushik Dey ·

    ASALT: Adaptive State Alignment for Lateral Transfer in Multi-agent Reinforcement Learning

    arXiv:2606.24601v1 Announce Type: new Abstract: Multi-agent reinforcement learning (MARL) addresses the problem of training multiple agents that pursue collaborative, competitive, or mixed objectives. Prior work has investigated transfer learning between source and target domains…

  268. arXiv cs.AI TIER_1 English(EN) · Bingnan Xiao, Chenhao Yang, Wei Ni, Xin Wang, Tony Q. S. Quek ·

    Agentic AI for Bilevel Long-Term Optimization of Policy-Driven Physical Layer Systems

    arXiv:2606.24416v1 Announce Type: new Abstract: Network operators' changing policies, service requirements, and stringent real-time constraints render existing methods designed with fixed objectives and constraints ineffective. This paper presents Agentic long-term performance op…

  269. arXiv cs.LG TIER_1 English(EN) · Manuel Baltieri ·

    Compositional Behavioral Semantics for State Abstraction in Reinforcement Learning

    State abstraction plays a key role in scaling reinforcement learning to complex but structured systems. In studying such systems, a wide range of behavioral structures have been studied in reinforcement learning, including value functions, invariants, bisimulation relations, and …

  270. Hugging Face Daily Papers TIER_1 English(EN) ·

    Stagnant Neuron: Towards Understanding the Plasticity Loss in Multi-Agent Reinforcement Learning Value Factorization Methods

    Multi-Agent Reinforcement Learning (MARL) value factorization methods can suffer from a loss of plasticity, gradually failing to adapt when transferring to new task instances. We trace this issue to stagnant neurons, units whose gradient updates become negligibly small relative t…

  271. arXiv cs.LG TIER_1 English(EN) · Cheng Wang ·

    Stagnant Neuron: Towards Understanding the Plasticity Loss in Multi-Agent Reinforcement Learning Value Factorization Methods

    Multi-Agent Reinforcement Learning (MARL) value factorization methods can suffer from a loss of plasticity, gradually failing to adapt when transferring to new task instances. We trace this issue to stagnant neurons, units whose gradient updates become negligibly small relative t…

  272. Hugging Face Daily Papers TIER_1 English(EN) ·

    Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It

    Research investigates how different supervisory signals and training strategies improve the stability and performance of large language models in tool-use tasks, addressing issues like catastrophic collapse and format sensitivity through interleaved supervised fine-tuning and rei…

  273. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Kaushik Dey ·

    GCT-MARL: Graph-Based Contrastive Transfer for Sample-Efficient Cooperative Multi-Agent Reinforcement Learning

    In cooperative multi-agent reinforcement learning (MARL), from a deployment perspective, it is challenging and expensive to train agents from scratch for each new environment or task. In this work, we propose GCT-MARL, a transfer learning framework that builds on the multi-view g…

  274. arXiv cs.AI TIER_1 English(EN) · Kaushik Dey ·

    ASALT: Adaptive State Alignment for Lateral Transfer in Multi-agent Reinforcement Learning

    Multi-agent reinforcement learning (MARL) addresses the problem of training multiple agents that pursue collaborative, competitive, or mixed objectives. Prior work has investigated transfer learning between source and target domains in MARL; however, the majority of existing appr…

  275. arXiv cs.AI TIER_1 English(EN) · Tony Q. S. Quek ·

    Agentic AI for Bilevel Long-Term Optimization of Policy-Driven Physical Layer Systems

    Network operators' changing policies, service requirements, and stringent real-time constraints render existing methods designed with fixed objectives and constraints ineffective. This paper presents Agentic long-term performance optimization (Agentic-LTPO), a nested bilevel opti…

  276. Hugging Face Daily Papers TIER_1 English(EN) ·

    Supervised Reinforcement Learning for the Coordination of Distributed Energy Resources

    The increasing integration of distributed energy resources (DERs) is crucial for power system decarbonization, yet unlocking DERs' flexibility is challenged by their inherent uncertainties and modelling complexity. As traditional optimization methods struggle with such uncertaint…

  277. Hugging Face Daily Papers TIER_1 English(EN) ·

    An Introduction to Causal Reinforcement Learning

    Causal inference provides a set of principles and tools that allow one to combine data and knowledge about an environment to reason with questions of counterfactual nature, i.e., what would have happened had reality been different, even when no data of this unrealized reality is …

  278. arXiv cs.CL TIER_1 English(EN) · Karan Singhal ·

    Reinforcement Learning Towards Broadly and Persistently Beneficial Models

    As AI systems are deployed across increasingly diverse and high-stakes settings, model alignment must generalize beyond the tasks and domains seen during training. This is especially important for reinforcement learning (RL), which can introduce unexpected misalignment through re…

  279. arXiv cs.CL TIER_1 English(EN) · Qi Zhang ·

    Group-Graph Policy Optimization for Long-Horizon Agentic Reinforcement Learning

    Group-based Reinforcement Learning (RL) has significantly enhanced Large Language Models (LLMs) in agentic scenarios. To achieve finer-grained policy updates, recent agentic RL frameworks have shifted from trajectory-level to step-level training. However, long-horizon agentic RL …

  280. arXiv cs.CL TIER_1 English(EN) · SingGuard Team ·

    SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning

    Vision-language models (VLMs) are increasingly deployed in consumer, medical, financial, and enterprise applications. This broad deployment expands the safety surface: risks can arise from multimodal question answering, assistant responses, and cross-modal composition, while mode…

  281. Hugging Face Daily Papers TIER_1 English(EN) ·

    Select-to-Act: Hierarchical Reinforcement Learning via Adaptive Language Guidance

    Reinforcement Learning (RL) has been widely applied to sequential decision-making, yet it often suffers from poor sample efficiency due to costly interactions with the environment. A limited line of recent work has started exploring improving RL efficiency by leveraging external …

  282. arXiv cs.AI TIER_1 English(EN) · Yanxi Chen, Weijie Shi, Yuexiang Xie, Boyi Hu, Yaliang Li, Bolin Ding, Jingren Zhou ·

    Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning

    arXiv:2606.20002v1 Announce Type: cross Abstract: This work presents a general framework for training large language models (LLMs) to "Connect the Dots" (CoD), a meta-capability required by long-lifecycle agents: as an LLM-based AI agent gets deployed in an environment, it solves…

  283. arXiv cs.AI TIER_1 English(EN) · Jingren Zhou ·

    Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning

    This work presents a general framework for training large language models (LLMs) to "Connect the Dots" (CoD), a meta-capability required by long-lifecycle agents: as an LLM-based AI agent gets deployed in an environment, it solves a long sequence of tasks while continuously explo…

  284. Hugging Face Daily Papers TIER_1 English(EN) ·

    Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning

    Large language models can be trained through reinforcement learning to develop a meta-capability enabling continuous learning and adaptation across long sequences of tasks in dynamic environments.

  285. arXiv stat.ML TIER_1 English(EN) · Sara Rajaram, R. James Cotton, Fabian H. Sinz ·

    Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning

    arXiv:2506.12529v2 Announce Type: replace-cross Abstract: Preference-based Reinforcement Learning (PbRL) entails a variety of approaches for aligning models with human intent to alleviate the burden of reward engineering. However, most previous PbRL work has not investigated the …

  286. arXiv stat.ML TIER_1 English(EN) · Hari Prasad ·

    Auditing the Risk Claims of Distributional Reinforcement Learning

    arXiv:2607.11607v1 Announce Type: cross Abstract: Distributional reinforcement learning agents learn full return distributions that are increasingly read at face value: for interpretability, risk-sensitive control, and safety monitoring. We ask a question theory anticipates but t…

  287. arXiv stat.ML TIER_1 English(EN) · Wen-Ting Wang ·

    Reinforcement Learning for Execution under Dynamic Fees in a Closed-Loop DEX Simulator

    arXiv:2607.10960v1 Announce Type: cross Abstract: Trader-facing dynamic fees are increasingly proposed for automated market makers (AMMs), but historical data do not identify how order flow would respond: trader-facing fees do not vary, trader types are latent, and a replayed tap…

  288. arXiv stat.ML TIER_1 English(EN) · Asim H. Gazi, Yongyi Guo, Daiqi Gao, Ziping Xu, Kelly W. Zhang, Susan A. Murphy ·

    Reinforcement Learning in the Real World: A Survey of Statistical Challenges and Future Directions

    arXiv:2601.15353v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has achieved remarkable success in real-world decision-making across diverse domains, including gaming, robotics, online advertising, public health, and natural language processing. Despite thes…

  289. arXiv stat.ML TIER_1 English(EN) · Yang Xu, Swetha Ganesh, Vaneet Aggarwal ·

    Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning

    arXiv:2506.07040v4 Announce Type: replace-cross Abstract: We study model-free methods for distributionally robust infinite-horizon average-reward Markov decision processes (MDPs). We present non-asymptotic convergence analyses of Q-learning and actor-critic algorithms for robust …

  290. arXiv stat.ML TIER_1 English(EN) · Miao Lu, Han Zhong, Tong Zhang, Jose Blanchet ·

    Distributionally Robust Reinforcement Learning with Interactive Data Collection: Fundamental Hardness and Near-Optimal Algorithms

    arXiv:2404.03578v3 Announce Type: replace-cross Abstract: The sim-to-real gap, which represents the disparity between training and testing environments, poses a significant challenge in reinforcement learning (RL). A promising approach to addressing this challenge is distribution…

  291. arXiv stat.ML TIER_1 English(EN) · Hari Prasad ·

    Auditing the Risk Claims of Distributional Reinforcement Learning

    Distributional reinforcement learning agents learn full return distributions that are increasingly read at face value: for interpretability, risk-sensitive control, and safety monitoring. We ask a question theory anticipates but that has not been measured directly: are the risk c…

  292. arXiv stat.ML TIER_1 English(EN) · Wen-Ting Wang ·

    Reinforcement Learning for Execution under Dynamic Fees in a Closed-Loop DEX Simulator

    Trader-facing dynamic fees are increasingly proposed for automated market makers (AMMs), but historical data do not identify how order flow would respond: trader-facing fees do not vary, trader types are latent, and a replayed tape is not a sequential decision environment. We the…

  293. arXiv stat.ML TIER_1 English(EN) · Zijie Cheng, Yang Peng, Zhihua Zhang ·

    Statistical Efficiency and Inference of Quantile Distributional Reinforcement Learning

    arXiv:2607.08444v1 Announce Type: new Abstract: In this paper, we study quantile-based distributional reinforcement learning from the perspective of statistical efficiency. We focus on distributional policy evaluation, whose goal is to characterize the return distribution, namely…

  294. arXiv stat.ML TIER_1 English(EN) · Zhihua Zhang ·

    Statistical Efficiency and Inference of Quantile Distributional Reinforcement Learning

    In this paper, we study quantile-based distributional reinforcement learning from the perspective of statistical efficiency. We focus on distributional policy evaluation, whose goal is to characterize the return distribution, namely the distribution of discounted cumulative rewar…

  295. arXiv stat.ML TIER_1 English(EN) · Zhihua Zhang ·

    Statistical Efficiency and Inference of Quantile Distributional Reinforcement Learning

    In this paper, we study quantile-based distributional reinforcement learning from the perspective of statistical efficiency. We focus on distributional policy evaluation, whose goal is to characterize the return distribution, namely the distribution of discounted cumulative rewar…

  296. arXiv stat.ML TIER_1 English(EN) · Nikita Yudin ·

    Mathematical methods of reinforcement learning

    Reinforcement learning (RL) is increasingly grounded in tools from probability, optimization, and operator theory. This survey organizes the mathematical structures that underpin the design and analysis of modern algorithms in RL. We begin from Markov decision processes (MDPs) an…

  297. arXiv cs.CV TIER_1 English(EN) · Dengyang Jiang, Dongyang Liu, Zanyi Wang, Qilong Wu, Liuzhuozheng Li, Hengzhuang Li, Xin Jin, David Liu, Changsheng Lu, Zhen Li, Bo Zhang, Mengmeng Wang, Steven Hoi, Peng Gao, Harry Yang ·

    Distribution Matching Distillation Meets Reinforcement Learning

    arXiv:2511.13649v5 Announce Type: replace Abstract: Distribution Matching Distillation (DMD) facilitates efficient inference by distilling multi-step diffusion models into few-step variants. Concurrently, Reinforcement Learning (RL) has emerged as a vital tool for aligning genera…

  298. arXiv stat.ML TIER_1 English(EN) · Stefano Masini, Cecilia Viscardi, Michela Baccini ·

    Full Bayesian Reinforcement Learning via LF-IBIS

    arXiv:2607.01741v1 Announce Type: new Abstract: Reinforcement Learning (RL) is a sequential decision-making framework in which an agent learns optimal policies through interaction with an environment by maximizing cumulative rewards. Among RL methods, Bayesian Reinforcement Learn…

  299. arXiv stat.ML TIER_1 English(EN) · Michela Baccini ·

    Full Bayesian Reinforcement Learning via LF-IBIS

    Reinforcement Learning (RL) is a sequential decision-making framework in which an agent learns optimal policies through interaction with an environment by maximizing cumulative rewards. Among RL methods, Bayesian Reinforcement Learning (BRL) addresses common practical challenges …

  300. arXiv stat.ML TIER_1 English(EN) · Yingru Li, Jiacai Liu, Jiawei Xu, Yuxuan Tong, Ziniu Li, Qian Liu, Baoxiang Wang ·

    Trust Region Masking for Long-Horizon LLM Reinforcement Learning

    arXiv:2512.23075v5 Announce Type: replace-cross Abstract: Policy gradient methods for Large Language Models optimize a policy $\pi_\theta$ via a surrogate objective computed from samples of a rollout policy $\pi_{\text{roll}}$. However, modern LLM-RL pipelines suffer from unavoid…

  301. Towards AI TIER_1 English(EN) · Fousseyni Sangaré ·

    Reinforcement Learning: Application of Value Iteration in Robotics Navigation Task

    <p>Hello everyone 😀 Welcome to this blog post !</p><p>Today we are going to see a real world application of Dynamic Programming in a Robotics Task. In a previous <a href="https://medium.com/@fousseyni.phd/rl-in-discrete-world-dynamic-programming-part2-generalized-policy-iteration…

  302. Towards AI TIER_1 English(EN) · Michael Ariaga ·

    Confidence Aware Reinforcement Learning: Advancing Large Language Models in Dynamic Environments

    <h4>Building Large Language Model Predictive Confidence to Navigate Uncertainty with Resiliency and Conviction</h4><p>Environments are in constant change as the physical world and contextual signals evolve to reflect new meaning or redefine ground truth. Large language models (LL…

  303. r/LocalLLaMA TIER_1 English(EN) · /u/de4dee ·

    [2607.07508] Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1uw7rm8/260707508_singlerollout_asynchronous_optimization/"> <img alt="[2607.07508] Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning" src="https://external-preview.redd.it/q3evP6JeDp…

  304. Mastodon — mastodon.social TIER_1 English(EN) · strike007 ·

    Moving beyond heuristic safety, non-vacuous generalization bounds are critical for formal verification in reinforcement learning. If we cannot mathematically gu

    Moving beyond heuristic safety, non-vacuous generalization bounds are critical for formal verification in reinforcement learning. If we cannot mathematically guarantee that an agent stays within its policy constraints, deploying it in high-stakes environments is reckless. # AI # …