PulseAugur
EN
LIVE 08:05:03

AI research advances reinforcement learning for math, adaptation, and fairness · 10 sources tracked

Researchers are exploring advanced reinforcement learning techniques to improve AI's mathematical reasoning and adaptation capabilities. One paper introduces Function-Structured Graph Reinforcement Learning (FSG-RL) to enhance mathematical problem-solving by generating code from function graphs and using multi-verifier feedback, significantly boosting accuracy. Other studies focus on optimizing Proximal Policy Optimization (PPO) through sample reuse and exploring data-free methods for test-time adaptation. Additionally, new approaches are being developed for unsupervised reinforcement learning, continual learning with neuroevolution, and integrating fairness with explainability in AI systems. AI

IMPACT Advances in reinforcement learning techniques could lead to more capable AI for complex reasoning, adaptation, and ethical considerations.

RANK_REASON Multiple research papers published on arXiv detailing novel methods and analyses in reinforcement learning.

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 149 sources. How we write summaries →

AI research advances reinforcement learning for math, adaptation, and fairness · 10 sources tracked

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
Multiple research papers published on arXiv detailing novel methods and analyses in reinforcement learning.
Source corroboration
149 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
13 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+65 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

Full methodology in our editorial standards.

COVERAGE [149]

  1. arXiv cs.LG TIER_1 English(EN) · Ocan Sankur (DEVINE), Thierry J\'eron (DEVINE), Nicolas Markey (DEVINE), David Mentr\'e (MERCE-France), Reiya Noguchi ·

    Requirement-Based Testing: Enhancing Reinforcement Learning with Game Theory

    arXiv:2407.18994v2 Announce Type: replace-cross Abstract: We consider the automatic online synthesis of black-box test cases from functional requirements specified as automata for reactive implementations. The goal of the tester is to reach some given state, so as to satisfy a co…

  2. arXiv cs.CL TIER_1 English(EN) · Zirui Li, Rech Silas, Lauri Juvela, Tom Backstrom, Mikko Kurimo ·

    EmphTTS: an emphasis-control TTS with reinforcement learning

    arXiv:2609.27599v2 Announce Type: cross Abstract: Generating controllable and human-like emphasis remains an open challenge in text-to-speech, even when explicit emphasis control signals are provided in the text input, limiting the communicative accuracy of synthetic speech in re…

  3. arXiv cs.CL TIER_1 English(EN) · Yifan Zhang ·

    On KL-Regularized Policy Optimization

    arXiv:2610.08963v1 Announce Type: cross Abstract: Asynchronous reinforcement learning (RL) for large language model (LLM) agents trains one policy on trajectories generated by another: rollouts come from stale checkpoints, and the inference engine's probabilities differ from the …

  4. arXiv cs.CL TIER_1 English(EN) · Shaokun Zhang, Yifan Zhang, Jian Hu, Yueying Li, Hao Zhang, Binfeng Xu, Jan Kautz, Yi Dong ·

    LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches

    arXiv:2610.06647v2 Announce Type: replace Abstract: Reinforcement learning has greatly advanced the capabilities of large language models, but its memory demands remain a barrier to broader adoption. We introduce LoGRA, an approach to RL post-training that reduces memory by retai…

  5. arXiv cs.LG TIER_1 English(EN) · Shripad V. Deshmukh, Yaswanth Chittepu, Dhawal Gupta, Philip Thomas, Scott Niekum ·

    Convex-Concave Reinforcement Learning

    arXiv:2610.09108v1 Announce Type: new Abstract: Policy learning drives many of the most consequential and heavily-invested applications of reinforcement learning today. Yet the core optimization problem it rests on (maximizing expected return) is notoriously non-convex, even unde…

  6. arXiv cs.LG TIER_1 English(EN) · Michael Tang, Mahmoud Abdelgalil, Jorge I. Poveda ·

    The Deceptive Bandit Problem: Exploratory Coupling and the Fragility of Multi-Agent Learning

    arXiv:2610.09120v1 Announce Type: new Abstract: Randomized exploration is central to bandit learning, multi-agent reinforcement learning, and zeroth-order policy search, yet its independence and privacy are usually only treated as technical assumptions. We show that these propert…

  7. arXiv cs.LG TIER_1 English(EN) · John L. Zhou, Yuxuan Dong, Jonathan C. Kao ·

    An Informational Curse of Horizon in Goal-Conditioned Policy Learning

    arXiv:2610.09247v1 Announce Type: new Abstract: The difficulty of learning goal-reaching policies is often attributed to a "curse of horizon" that manifests as bias accumulation in temporal-difference backups and noisy advantage estimates. In this work, we identify an additional …

  8. arXiv cs.LG TIER_1 English(EN) · Zicheng Hu, Zhijian Zhou, Xuan Zhang, Yuchen Liu, Cheng Chen, Yuan Li, Qi Gu, Yan Feng, Hongyan Hao, Chao Qu ·

    COPC: Coupled Off-Policy Correction for Asynchronous LLM Reinforcement Learning

    arXiv:2610.09597v1 Announce Type: new Abstract: Asynchronous RL accelerates large language model post-training by decoupling rollout generation from optimization, but trains on stale trajectories. Existing methods primarily correct token-level policy mismatch through importance-r…

  9. arXiv cs.LG TIER_1 English(EN) · Mahmoud Selim, Cristina Cipriani, Karl Henrik Johansson ·

    Beyond Policy Support: Interaction Constrained Offline Reinforcement Learning for Autonomous Driving

    arXiv:2610.09763v1 Announce Type: new Abstract: Offline reinforcement learning enables reward-driven policy improvement from fixed datasets without requiring online exploration, making it particularly attractive in safety-critical domains. A central challenge, however, is distrib…

  10. arXiv cs.LG TIER_1 English(EN) · Yongqiang Yao, Jinru Tan, Kaihuan Liang, Zixin Yin, Yazhe Niu, Ruihao Gong, Dahua Lin, Ningyi Xu ·

    RollVerify: Bridging Efficiency and Accuracy in Long-Tail Rollout Reinforcement Learning

    arXiv:2610.09914v1 Announce Type: new Abstract: Reinforcement learning is crucial for improving large language models' reasoning and generalization. It relies on massive rollouts whose lengths become increasingly long-tailed as context windows grow. In on-policy training, these l…

  11. arXiv cs.LG TIER_1 English(EN) · Tommaso Marzi, Ahmed Hendawy, Jan Peters, Carlo D'Eramo, Andrea Cini, Cesare Alippi ·

    Continual Graph Multi-Agent Reinforcement Learning

    arXiv:2610.10302v1 Announce Type: new Abstract: In Continual Multi-Agent Reinforcement Learning (CMARL), agents learn cooperative policies across sequences of tasks, aiming to adapt effectively to new tasks while preserving the ability to solve previously encountered ones. In man…

  12. arXiv cs.LG TIER_1 English(EN) · Huizhen Yu, Isaiah Heidt ·

    Average-Reward Reinforcement Learning for Multichain MDPs: A Hierarchical Decomposition Approach

    arXiv:2610.10326v1 Announce Type: new Abstract: We study learning optimal policies in average-reward multichain Markov decision processes (MDPs), where the optimal gain may depend on the initial state and recurrence structures vary across policies, creating challenges for reinfor…

  13. arXiv cs.LG TIER_1 English(EN) · Fernando Martinez, Abhishek Satyam, Tao Li, Junaid Farooq, Ying Wang, Juntao Chen ·

    Ask the Expert: LLM-Guided Reinforcement Learning for Autonomous Cyber Defense

    arXiv:2610.09337v1 Announce Type: cross Abstract: Policy-based reinforcement learning (RL) approaches have produced promising results for autonomous cyber defense; however, they are sample-inefficient in settings where defenders must respond under delayed, partial observations wi…

  14. arXiv cs.LG TIER_1 English(EN) · Lijun Bo, Yijie Huang, Chenhao Lu ·

    Boundary-aware Reinforcement Learning for Hypercube State Spaces via Deterministic Policy Gradient

    arXiv:2610.09712v1 Announce Type: cross Abstract: We develop a continuous-time deterministic policy gradient framework for reinforcement learning with reflected state dynamics, where the state process is governed by a controlled reflected stochastic differential equation on a hyp…

  15. arXiv cs.LG TIER_1 English(EN) · Benedikt Koch, Winston Chou, Aur\'elien Bibaut, Nathan Kallus ·

    Policy Learning with Weak Signals

    arXiv:2610.10167v1 Announce Type: cross Abstract: Policy learning in digital experimentation faces three challenges: weak signal-to-noise ratios, rich covariate spaces, and massive data volumes. We formalize this regime by modeling treatment-effect estimates from increasingly fin…

  16. arXiv cs.LG TIER_1 English(EN) · Peng Xia, Jianwen Chen, Hanyang Wang, Jiaqi Liu, Kaide Zeng, Yu Wang, Siwei Han, Yiyang Zhou, Xujiang Zhao, Haifeng Chen, Zeyu Zheng, Cihang Xie, Huaxiu Yao ·

    SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning

    arXiv:2602.08234v2 Announce Type: replace Abstract: Large Language Model (LLM) agents have shown stunning results in complex tasks, yet they often operate in isolation, failing to learn from past experiences. Existing memory-based methods primarily store raw trajectories, which a…

  17. arXiv cs.AI TIER_1 English(EN) · Ange Tong ·

    Learning When to Refine: Long-Horizon Reinforcement Learning for Budgeted Neural-Operator PDE Solvers

    arXiv:2610.06883v1 Announce Type: cross Abstract: Neural operators provide fast surrogates for time-dependent PDEs, but autoregressive deployment creates a refinement-allocation problem: prediction errors vary over space and time, while only a finite number of local corrections c…

  18. arXiv cs.LG TIER_1 English(EN) · Silong Yong, Stephen Sheng, Carl Qi, Xiaojie Wang, Evan Sheehan, Anurag Shivaprasad, Yaqi Xie, Katia Sycara, Yesh Dattatreya ·

    Generalizable Dense Reward for Long-Horizon Robotic Tasks

    arXiv:2604.00055v2 Announce Type: replace-cross Abstract: Existing robotic foundation policies are trained primarily via large-scale imitation learning. While such models demonstrate strong capabilities, they often struggle with long-horizon tasks due to distribution shift and er…

  19. arXiv cs.LG TIER_1 English(EN) · Valliappan Chidambaram Adaikkappan, Sai Rajeswar, Pietro Mazzaglia, David Meger ·

    MSPR: Multi-scale Predictive Representations for Goal-conditioned Reinforcement Learning

    arXiv:2605.09364v2 Announce Type: replace Abstract: This paper investigates robust representation learning in offline goal-conditioned reinforcement learning (GCRL). Particularly in sparse reward scenarios, learning representations that align state and goal latents is a challenge…

  20. arXiv cs.LG TIER_1 English(EN) · Junsoo Ha ·

    Stochastic Gradient Descent Ascent is Suboptimal for Nonconvex-PL Min-Max Games

    arXiv:2610.07814v1 Announce Type: cross Abstract: How far can stochastic gradient descent ascent (SGDA) go by tuning its timescale ratio and step sizes in nonconvex min-max games? We answer this question for nonconvex-PL (NC-PL) games by establishing the first tight complexity of…

  21. arXiv cs.LG TIER_1 English(EN) · Fernando Martinez, Tao Li, Yingdong Lu, Juntao Chen ·

    Independent Multi-Agent Reinforcement Learning with Counterfactual Semantic-Social World Models

    arXiv:2610.07704v1 Announce Type: cross Abstract: Fully decentralized multi-agent reinforcement learning (MARL), also referred to as independent learning, requires each agent to learn and act using only its local information and experience, without a centralized critic or inter-a…

  22. arXiv cs.LG TIER_1 English(EN) · Naoki Nishikawa, Taiji Suzuki ·

    Reinforcement Learning for Hierarchical Reasoning Rewards: Minimax-Optimal Rates with Transformers

    arXiv:2610.08561v1 Announce Type: new Abstract: Reinforcement learning (RL) has become a standard tool for post-training language models on reasoning tasks, where the policy is updated by reward feedback while exploring the space of responses. Despite its empirical success, theor…

  23. arXiv cs.LG TIER_1 English(EN) · Fengxu Liu, Siwei Wang, Gal Dalal, Shie Mannor, Yihan Du ·

    Reinforcement Learning with Segment Reward Feedback under Linear Function Approximation

    arXiv:2610.08271v1 Announce Type: new Abstract: Classical reinforcement learning (RL) assumes that a reward is observed for every visited state-action pair. However, in real-world applications such as autonomous driving, such fine-grained feedback can be costly or difficult to co…

  24. arXiv cs.LG TIER_1 English(EN) · SungJae Ahn, Jeong Woon Lee, Kyoleen Kwak, Hyoseok Hwang ·

    Revisiting Temporal Regularization for Smooth Control in Deep Reinforcement Learning

    arXiv:2610.07910v1 Announce Type: new Abstract: Deep Reinforcement Learning policies can produce nonsmooth action oscillations that hinder deployment on physical robots. Existing architectural and penalty-based approaches seek spatial smoothness by directly reducing sensitivity t…

  25. arXiv cs.CL TIER_1 English(EN) · Bolian Li, Ting-Yao Hu, Cheng-Yu Hsieh, Sanjoy Chowdhury, Oncel Tuzel, Raviteja Vemulapalli ·

    Structuring MoE Expert Selection for Agentic Reinforcement Learning

    arXiv:2610.07332v1 Announce Type: cross Abstract: Long-horizon LLM agents are frequently implemented using sparse mixture-of-experts (MoE) models, yet the co-design of agentic behavior and MoE structures remains underexplored. In this work, we comprehensively study the connection…

  26. arXiv cs.AI TIER_1 English(EN) · Chenliang Li, Adel Elmahdy, Alex Boyd, Zhongruo Wang, Siliang Zeng, Alfredo Garcia, Parminder Bhatia, Taha Kass-Hout, Cao Xiao, Mingyi Hong ·

    Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization

    arXiv:2511.20718v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) algorithms such as PPO and GRPO are widely used to train large language models (LLMs) for multi-turn agentic tasks. However, in off-policy training pipelines, these methods can exhibit unstable …

  27. arXiv cs.AI TIER_1 English(EN) · Wenwen Si, Honghao Wei ·

    Reinforcement Learning with Conformal Action Sets: An Application to Sequential Recommendation

    arXiv:2610.08743v1 Announce Type: cross Abstract: Sequential recommenders typically use a fixed slate size even though the number of useful alternatives changes within a session. We propose Reinforcement Learning with Calibrated Pruning (RLCP), which adapts the retained action se…

  28. arXiv cs.AI TIER_1 English(EN) · Guhyeon Kang, Minhae Kwon ·

    Variance-Averse $n$-Step Offline Reinforcement Learning for Sparse Long-Horizon Environments

    arXiv:2610.07899v1 Announce Type: cross Abstract: Generative actors are transforming offline reinforcement learning (RL) by enabling expressive policy classes that model complex action distributions. However, this expressiveness also exposes a key challenge in heterogeneous datas…

  29. arXiv cs.AI TIER_1 English(EN) · Qi Liu, Fengming Liang, Yiqun Chen, Erhan Zhang, Jiaxin Mao ·

    Learning to Retrieve via Reinforcement Learning in Embedding Space

    arXiv:2610.07731v1 Announce Type: cross Abstract: Dense retrieval models are typically trained with contrastive objectives that learn effective representations but do not directly optimize retrieval metrics or downstream task performance. To address this problem, we introduce REL…

  30. arXiv cs.AI TIER_1 English(EN) · Xiaoyang Cao, Jingqi Li, Zhe Fu, Alexandre M. Bayen ·

    Who Bears the Burden? Learning Responsibility for Shared Constraints in Multi-Agent Reinforcement Learning

    arXiv:2610.07491v1 Announce Type: cross Abstract: When multiple agents share a cost budget, a common Lagrange multiplier can enforce the aggregate constraint but does not determine how its penalty should be allocated across agents. Uniform penalties ignore heterogeneity in the re…

  31. arXiv cs.AI TIER_1 English(EN) · T. Y. Tsui, Zihao Ye, Pengxiang Cai, Yanchao Li, Yuqiang Li, Zhehong Ai ·

    Minimal Witness Reinforcement Learning

    arXiv:2610.07226v1 Announce Type: cross Abstract: ``What are the irreducible conditions that are sufficient to produce an outcome?'' is one of the most common questions that recur across computation and science. Its answers, the minimal sufficient witnesses, are what we mean by e…

  32. arXiv cs.AI TIER_1 English(EN) · Zhen Li, Shuai Zhang, Yanggan Gu, Yiming Zhang, Yang Yu, Mingfa Feng, Congkai Xie, Shuang Yu, Junjie Lai, Hongxia Yang ·

    TRIAGE: Direction-Aware Mismatch Stabilization of Native NVFP4 Reinforcement Learning

    arXiv:2610.07043v1 Announce Type: cross Abstract: Low-precision execution can substantially accelerate reinforcement learning (RL) for large language models, but discrepancies between learner and sampler execution can destabilize policy optimization. In this paper, we characteriz…

  33. Hugging Face Daily Papers TIER_1 English(EN) ·

    Revisiting Temporal Regularization for Smooth Control in Deep Reinforcement Learning

    Deep Reinforcement Learning policies can produce nonsmooth action oscillations that hinder deployment on physical robots. Existing architectural and penalty-based approaches seek spatial smoothness by directly reducing sensitivity to changes in state inputs, but their broad const…

  34. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Jiaxin Mao ·

    Learning to Retrieve via Reinforcement Learning in Embedding Space

    Dense retrieval models are typically trained with contrastive objectives that learn effective representations but do not directly optimize retrieval metrics or downstream task performance. To address this problem, we introduce RELER (REinforcement LEarning for Retrieval), a reinf…

  35. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Juntao Chen ·

    Independent Multi-Agent Reinforcement Learning with Counterfactual Semantic-Social World Models

    Fully decentralized multi-agent reinforcement learning (MARL), also referred to as independent learning, requires each agent to learn and act using only its local information and experience, without a centralized critic or inter-agent communication. Such a stringent information s…

  36. Hugging Face Daily Papers TIER_1 English(EN) ·

    On KL-Regularized Policy Optimization

    Asynchronous reinforcement learning (RL) for large language model (LLM) agents trains one policy on trajectories generated by another: rollouts come from stale checkpoints, and the inference engine's probabilities differ from the trainer's even at identical parameters. Standard r…

  37. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Alexandre M. Bayen ·

    Who Bears the Burden? Learning Responsibility for Shared Constraints in Multi-Agent Reinforcement Learning

    When multiple agents share a cost budget, a common Lagrange multiplier can enforce the aggregate constraint but does not determine how its penalty should be allocated across agents. Uniform penalties ignore heterogeneity in the rewards agents sacrifice, while agent-specific multi…

  38. Hugging Face Daily Papers TIER_1 English(EN) ·

    ThunderSyncRL: Lossless Acceleration of Agentic Reinforcement Learning

    Language models are moving beyond generating answers to pursuing long-horizon goals in interactive environments. Post-training these agents requires long, heterogeneous trajectories, and synchronous systems leave learner engines idle until rollout and verification finish. To sque…

  39. arXiv cs.LG TIER_1 Deutsch(DE) · Jinwoo Kim, Shraddha Barke ·

    Understanding Enrichment in Reinforcement Learning

    arXiv:2610.02846v1 Announce Type: new Abstract: When rewards are sparse, reinforcement learning with verifiable rewards (RLVR) often uses hints or intermediate guidance to generate more successful rollouts. This enrichment biases policy-gradient updates unless corrected via impor…

  40. arXiv cs.CL TIER_1 English(EN) · Xin Wang, Wenhao Wu, Menghao Zhang, Zhi Wang, Kun Shao, Jian Luan ·

    AdaStep: Adaptive Step Credit Weighting for Agentic Reinforcement Learning

    arXiv:2610.03223v1 Announce Type: cross Abstract: Long-horizon LLM agents are typically trained with sparse outcome rewards, making trajectory-level objectives too coarse to distinguish the contribution of individual decisions. Step-level credit assignment provides finer-grained …

  41. arXiv cs.LG TIER_1 English(EN) · Ganghun Lee, Minji Kim, Minsu Lee, Byoung-Tak Zhang ·

    Reward Inflation: A Healthy Stimulus for Reinforcement Learning

    arXiv:2610.02545v1 Announce Type: new Abstract: Reward serves as the primary learning signal in reinforcement learning (RL). However, while reward magnitudes are typically held fixed throughout training, their temporal modulation remains underexplored. In this paper, we propose r…

  42. arXiv cs.AI TIER_1 English(EN) · Arip Asadulaev, Aladin Djuhera, Karim Salta, Holger Boche, Fakhri Karray, Martin Takac ·

    Tropical Reinforcement Learning

    arXiv:2610.02478v1 Announce Type: new Abstract: Reinforcement learning for large language models typically maximizes expected return, adding up the probabilities of all successful trajectories. However, the classical sum formulation can only report how often the model policy succ…

  43. arXiv cs.AI TIER_1 English(EN) · Mingyu Park, Samyeul Noh, Hyun Myung, Donghwan Lee ·

    VIGOR: Zero-Shot Visual Generalization via Latent-Space Consistency in Model-Based Reinforcement Learning

    arXiv:2610.02801v1 Announce Type: new Abstract: Model-based reinforcement learning (MBRL) achieves strong sample efficiency by planning within learned latent dynamics, yet its performance degrades substantially under unseen visual distractions such as background variations, light…

  44. arXiv cs.AI TIER_1 English(EN) · Yuxuan Fan, Jaehong Yoon ·

    MetaRubric: Learning to Reward for Rubric-Based Reinforcement Learning

    arXiv:2610.02824v1 Announce Type: new Abstract: Rubric-based reinforcement learning extends reward-driven optimization to open-ended tasks by assigning partial credit to individual response requirements. However, rubric judges can assign a high criterion score even when the infor…

  45. arXiv cs.AI TIER_1 English(EN) · Xinjie Liu, Ruihan Zhao, Anirban Chaudhuri, Cyrus Neary, Ufuk Topcu, David Fridovich-Keil ·

    Multi-Fidelity Policy Gradients Stabilize Data-Scarce Reinforcement Learning

    arXiv:2610.02505v1 Announce Type: cross Abstract: Policy gradient methods for on-policy reinforcement learning (RL) can become unstable when expensive, scarce target-domain data yield noisy gradient estimates. We address this challenge by complementing limited high-fidelity (HF) …

  46. arXiv cs.LG TIER_1 English(EN) · Juri Pfammatter, Kaixian Qu, Clemens Schwarke, Victor Klemm, Marco Hutter ·

    Bidirectional Voronoi-biased Exploration Curriculum for Reinforcement Learning

    arXiv:2610.03395v1 Announce Type: new Abstract: Long-horizon tasks with sparse rewards pose an exploration bottleneck for goal-conditioned reinforcement learning: a policy started from the initial state rarely reaches the goal and receives no learning signal. Reference motions, h…

  47. arXiv cs.LG TIER_1 English(EN) · Amit Thakur, Mukesh Singhal ·

    Turnover-Orthogonal Credit Assignment for Open-Team Multi-Agent Reinforcement Learning

    arXiv:2610.02847v1 Announce Type: cross Abstract: Open-team multi-agent reinforcement learning studies cooperative systems in which agents may join, leave, or be replaced during an episode. In such settings, the team return changes both because agents choose useful actions and be…

  48. arXiv cs.LG TIER_1 English(EN) · Matthew Brun, Xu Andy Sun ·

    On the Convergence of Success Conditioning for Policy Optimization

    arXiv:2610.03642v1 Announce Type: new Abstract: Success conditioning is a strategy for improving decision-making policies in stochastic environments; it updates a policy by increasing the probability of taking actions that yield successful outcomes. Success conditioning is common…

  49. arXiv cs.LG TIER_1 English(EN) · Yudong Lin, Haoyuan Deng, Zhuoxuan Yuan, Zaijia Yang, Yuanjiang Xue, Ziwei Wang ·

    UniIntervene++: An Adaptive Intervention Agent for Efficient Real-World Reinforcement Learning

    arXiv:2610.03620v1 Announce Type: new Abstract: Online reinforcement learning (RL) enables robot policies to improve through physical interaction, but the assistance they require changes as their competence evolves. Existing intervention strategies based on offline estimates or f…

  50. arXiv cs.AI TIER_1 English(EN) · Hongqiang Lin, Dongxu Zhang, Yiding Sun, Mingzhe Li, Ning Yang, Haijun Zhang ·

    Offline Policy Optimization with Posterior Sampling

    arXiv:2605.07393v2 Announce Type: replace Abstract: A fundamental challenge in model-based offline reinforcement learning (RL) lies in the trade-off between generalization and robustness against exploitation errors in out-of-distribution (OOD) regions. The key to resolving this t…

  51. arXiv cs.AI TIER_1 English(EN) · Jiaxin Zhang, Xiangyu Peng, Qinglin Chen, Yu Li, Hiroaki Hayashi, Chien-Sheng Wu ·

    Prospective Hindsight: Self-Calibrating Reinforcement Learning via Prediction-Reality Gaps

    arXiv:2610.02740v1 Announce Type: cross Abstract: Reinforcement learning for long-horizon agents relies on purely retrospective training signals: credit is assigned only after observing environmental consequences, leaving the agent's belief at action time invisible to the gradien…

  52. arXiv cs.AI TIER_1 English(EN) · Saptarshi Nath, Inish M. D'Souza, Antonio Carta, Soheil Kolouri, Andrea Soltoggio ·

    How to Find and Reuse Policies for Continuous Adaptation in Lifelong Reinforcement Learning

    arXiv:2610.03119v1 Announce Type: cross Abstract: In lifelong reinforcement learning, retaining previously learned policies is not sufficient for effective transfer to a new task. Useful knowledge may be distributed across several prior policies, and its relevance may change as t…

  53. arXiv cs.AI TIER_1 English(EN) · Guilhem Loussouarn, Nancy Nayak, Kin K. Leung ·

    Single or Multiple Policies for Phase-Structured Reinforcement Learning?

    arXiv:2610.03475v1 Announce Type: cross Abstract: Many reinforcement-learning (RL) problems are non-stationary yet structured and can be decomposed into phases, each with its own transition probabilities and reward functions. When the phase sequence is known, the common solution …

  54. arXiv cs.AI TIER_1 English(EN) · Jiaxuan Luo, Xingguo Xu, Shanshan Wang, Yuhan Zhou, Zhen Zhang ·

    CriticHack: Evaluating Visual Rewards Under Robot Policy Optimization

    arXiv:2610.02527v1 Announce Type: cross Abstract: Learned visual reward models are increasingly used to optimize robot policies, yet a reward model can score an execution that acts on the wrong object as highly as one that completes the task. We show that optimizing such a reward…

  55. Hugging Face Daily Papers TIER_1 English(EN) ·

    Structuring MoE Expert Selection for Agentic Reinforcement Learning

    Long-horizon LLM agents are frequently implemented using sparse mixture-of-experts (MoE) models, yet the co-design of agentic behavior and MoE structures remains underexplored. In this work, we comprehensively study the connections between agentic post-training and MoE expert sel…

  56. Hugging Face Daily Papers TIER_1 English(EN) ·

    TRIAGE: Direction-Aware Mismatch Stabilization of Native NVFP4 Reinforcement Learning

    Low-precision execution can substantially accelerate reinforcement learning (RL) for large language models, but discrepancies between learner and sampler execution can destabilize policy optimization. In this paper, we characterize the interaction between mismatch and the policy-…

  57. Hugging Face Daily Papers TIER_1 English(EN) ·

    LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches

    Reinforcement learning (RL) has greatly advanced the capabilities of large language models (LLMs), but its memory demands remain a barrier to broader adoption. We introduce LoGRA, an approach to RL post-training that reduces memory by retaining useful learning signals in low-rank…

  58. Hugging Face Daily Papers TIER_1 English(EN) ·

    Minimal Witness Reinforcement Learning

    ``What are the irreducible conditions that are sufficient to produce an outcome?'' is one of the most common questions that recur across computation and science. Its answers, the minimal sufficient witnesses, are what we mean by explanations, mechanisms and reasons. These problem…

  59. arXiv cs.LG TIER_1 English(EN) · Morgan Byrd, Maks Sorokin, Robert Wright, Sehoon Ha ·

    Reward as Observation: Learning Reward-Based Policies for Rapid Adaptation

    arXiv:2610.00729v1 Announce Type: new Abstract: This paper explores a reward-based policy to achieve zero-shot transfer between source and target environments with completely different observation spaces. While humans can demonstrate impressive adaptation capabilities, deep neura…

  60. arXiv cs.LG TIER_1 English(EN) · Eleni Nisioti, Andrea Cossu, Kathrin Korte, Sebastian Risi ·

    Continual Reinforcement Learning with Neuroevolution

    arXiv:2610.01583v1 Announce Type: cross Abstract: Despite many studies about causes and remedies of plasticity loss in Reinforcement Learning (RL) under continual task changes, no RL method has yet consistently achieved a good balance between adaptation and forgetting. Here we tu…

  61. arXiv cs.LG TIER_1 English(EN) · Tristan Shah, Wooyoung Chung, Volodomyr Makarenko, Juan Wachs, Stas Tiomkin ·

    Bellman Meets Lyapunov: Unsupervised Reinforcement Learning via Mastering Chaos

    arXiv:2610.02012v1 Announce Type: new Abstract: Reinforcement learning (RL) is a powerful paradigm for training agents, yet its success rests on domain expertise of human engineers who design informative reward signals for every new task. Unsupervised RL aims to reduce this engin…

  62. arXiv cs.LG TIER_1 English(EN) · Zihan Liu, Xurong Xie ·

    Function-Structured Reinforcement Learning with Executable Verifiers for Mathematical Reasoning

    arXiv:2610.01729v1 Announce Type: new Abstract: Algorithmic mathematical reasoning requires reliable decomposition, computation, and aggregation. Final-answer rewards provide limited guidance on intermediate errors, while successful execution does not guarantee mathematical corre…

  63. arXiv cs.LG TIER_1 English(EN) · Ryunyi Lee, Kangjun Noh, Somin Kim, Heedong Kim, Kyungwoo Song ·

    Range-GRPO: Policy Optimization via Pairwise Relations among Reward Intervals

    arXiv:2610.01548v1 Announce Type: new Abstract: As the use of large language models (LLMs) expands, post-training has become increasingly important for adapting them to downstream tasks. However, obtaining reliable supervision remains costly, especially in domains without referen…

  64. arXiv cs.LG TIER_1 English(EN) · Alessandro Montenegro, Riccardo Venturelli, Marco Mussi, Matteo Papini, Alberto Maria Metelli ·

    Reusing Past Samples in Proximal Policy Optimization: When and How Does It Help?

    arXiv:2610.01399v1 Announce Type: new Abstract: Among on-policy deep reinforcement learning methods, Proximal Policy Optimization (PPO) has become the de facto standard, due to its consistently strong empirical performance across diverse application domains. However, on-policy me…

  65. arXiv cs.LG TIER_1 English(EN) · Zhanming Zhang, Vinoth Selvendran ·

    Sharpen Before You Adapt: Data-Free Entry-State Sharpening for Test-Time Reinforcement Learning

    arXiv:2610.00903v1 Announce Type: new Abstract: Test-time reinforcement learning (TTRL) adapts language models on unlabeled test problems using supervision derived from their own samples. This makes the checkpoint's \emph{entry state} consequential: a diffuse policy provides nois…

  66. arXiv cs.LG TIER_1 English(EN) · Bente Hinkenhuis, Seyed Sahand Mohammadi Ziabari, Ali Mohammed Mansoor Alsahag ·

    Integrating Fairness and Explainability in a Multiple Instance Reinforcement Learning System

    arXiv:2610.00035v1 Announce Type: new Abstract: Predicting student performance from educational interaction data requires models that are both accurate and sufficiently transparent to support meaningful intervention, while demographic information introduces an additional risk of …

  67. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Mukesh Singhal ·

    Turnover-Orthogonal Credit Assignment for Open-Team Multi-Agent Reinforcement Learning

    Open-team multi-agent reinforcement learning studies cooperative systems in which agents may join, leave, or be replaced during an episode. In such settings, the team return changes both because agents choose useful actions and because the active population itself changes. Standa…

  68. arXiv cs.AI TIER_1 English(EN) · Jinhwan Sul, Alex Oshin, Evangelos A. Theodorou ·

    Reinforcement Learning to Accelerate Primal-Dual Hybrid Gradient for Linear Programming

    arXiv:2610.01546v1 Announce Type: cross Abstract: Primal-dual hybrid gradient (PDHG) methods solve large-scale linear programs (LPs) using GPU-friendly matrix-vector products and projections, but their practical performance depends on coordinating algorithm parameters, accelerati…

  69. arXiv cs.AI TIER_1 English(EN) · Pedro Robles Dutenhefner, Dikshant Shehmar, Wagner Meira Jr., Marlos C. Machado ·

    Learning Multiple Timescales for Goal-Conditioned Reinforcement Learning

    arXiv:2610.00849v1 Announce Type: new Abstract: Existing approaches to offline goal-conditioned reinforcement learning (GCRL) struggle with long-horizon tasks. Discounting shrinks value differences between distant states until they fall below the function approximation error, lea…

  70. arXiv cs.AI TIER_1 English(EN) · Ziyi Chen, Yan Zhang, Jianhui Wei, Daoan Zhang, Zuozhu Liu ·

    Dependency-Aware Reward Shaping for Agentic Reinforcement Learning

    arXiv:2610.01207v1 Announce Type: new Abstract: When training large language models with reinforcement learning, terminal rewards provide little guidance about which steps matter. Common methods for assigning step credit overlook that work built on uncorrected mistakes is wasted …

  71. arXiv cs.AI TIER_1 English(EN) · Shiu-Hong Kao, Yubo Zhao, Zhenyu Tian, Pengzhan Sun, Yicong Li, Angela Yao ·

    Rethinking Probability-Based Reinforcement Learning From Posterior Concentration

    arXiv:2610.01458v1 Announce Type: new Abstract: Verifier-free reinforcement learning with probability-based rewards offers a promising way to train LLMs on general reasoning tasks where external verifiers are unavailable. Yet the reliability of these rewards, especially in long-h…

  72. arXiv cs.AI TIER_1 English(EN) · Jude Waide, Robert Lieck ·

    Decision Titan: Test-Time Training for Long-Term Memory in Offline Reinforcement Learning

    arXiv:2610.01513v1 Announce Type: new Abstract: Long-term dependencies remain a major challenge for sequential decision-making in the field of AI: RNNs suffer from vanishing gradients and the limited expressivity of vector-based hidden states, whilst Transformer-based models are …

  73. arXiv cs.AI TIER_1 English(EN) · Abid Mohamed Nadhir, Ahmad Al Hanbali, Beggas Mounir ·

    Homomorphic Advantage Operator: Stabilizing Reinforcement Learning Under Fully Homomorphic Encryption Constraints

    arXiv:2610.02074v1 Announce Type: new Abstract: Privacy-preserving machine learning presents significant deployment challenges on the cloud for intelligent systems with confidential data. Fully Homomorphic Encryption (FHE) offers a compelling solution for secure computation, pres…

  74. arXiv cs.AI TIER_1 English(EN) · Bo-Wen Zhang, Junwei He, Maoqi Liu, Feiran Li, Song-Lin Lv, Wentao Ma, Rongyi Lin, Shuhan Zhong, Lan-Zhe Guo ·

    T2SPO: Trajectory-to-Step Policy Optimization for Agentic Reinforcement Learning

    arXiv:2610.00388v1 Announce Type: cross Abstract: Reinforcement learning enables large language model (LLM) agents to learn multi-step behaviors through interaction with their environments. However, rewards in many interactive tasks reflect only the final outcome, providing limit…

  75. arXiv cs.AI TIER_1 English(EN) · Oleg Shchendrigin, Egor Cherepanov, Aleksandr I. Panov, Alexey K. Kovalev ·

    ALER: Adaptive Learnable Experience Rewriting for Reinforcement Learning

    arXiv:2610.00592v1 Announce Type: cross Abstract: In partially observable reinforcement learning (RL), a later observation can make stored information obsolete or change what it implies for the next decision. Memory architectures and benchmarks for RL mostly test retention, the a…

  76. arXiv cs.AI TIER_1 English(EN) · Christos Spyridon Koulouris, Carlo Campajola ·

    Beyond Supra-Competitive Outcomes: Collusive Behaviour in Deep Reinforcement Learning for Optimal Execution Games

    arXiv:2610.00619v1 Announce Type: cross Abstract: In this paper, we extend earlier findings of supra-competitive outcomes in optimal-execution games by identifying a learned punitive mechanism that deters deviations and provides behavioural evidence of collusion. We investigate t…

  77. arXiv cs.AI TIER_1 English(EN) · Yujie Zhu, Charles A. Hepburn, Matthew Thorpe, Giovanni Montana ·

    Optimal Transport Meets Reinforcement Learning: A Survey

    arXiv:2610.01413v1 Announce Type: cross Abstract: Reinforcement learning (RL) algorithms frequently compare probability distributions, such as state visitation distributions induced by policies and experts, action distributions from learned policies and offline datasets, or trans…

  78. arXiv cs.AI TIER_1 English(EN) · Yafei Zhang, Songshuo Lu, Sicong Liao, Zhi Chen, Yaohua Tang ·

    CARM: Cancellation-Aware Response Masking for LLM Reinforcement Learning

    arXiv:2610.02039v1 Announce Type: cross Abstract: Recent years have witnessed the rapid adoption of reinforcement learning (RL) in large language model (LLM) post-training, with substantial gains in mathematical reasoning and code generation. In practical systems, however, policy…

  79. arXiv cs.AI TIER_1 English(EN) · Sebastian Sanokowski, Alireza Sarmadi, Majid Khadiv ·

    FERPO: Forward Entropy-Regularized Policy Optimization

    arXiv:2610.02198v1 Announce Type: cross Abstract: Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic. However, critics are typically trained to predict returns, and accurate value pr…

  80. arXiv cs.AI TIER_1 English(EN) · Debosmita Bhaumik, Julian Togelius, Georgios N. Yannakakis, Ahmed Khalifa ·

    Learning Local Constraints for Reinforcement-Learned Content Generators

    arXiv:2605.13570v2 Announce Type: replace Abstract: Constraint-based game content generators that learn local constraints from existing content, such as Wave Function Collapse (WFC), can generate visually satisfying game levels but face challenges in optimizing global properties,…

  81. arXiv cs.AI TIER_1 English(EN) · Yunho Choi, Jongwon Lim, Woojin Ahn, Minjae Oh, Jeonghoon Shim, Yohan Jo ·

    Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States

    arXiv:2605.07579v3 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) for Large Reasoning Models rests on variance reduction, which requires both a reliable baseline and high prompt diversity within each training batch. This is especially…

  82. arXiv cs.LG TIER_1 English(EN) · Sathya Kamesh Bhethanabhotla, Efstratios Gavves, Andr\'e Biedenkapp ·

    Reinforcement Learning with Complex (valued) Memories

    arXiv:2609.38598v1 Announce Type: new Abstract: Partially observable environments pose a fundamental challenge in deep reinforcement learning, requiring agents to compress temporal information from observations and maintain a memory to make effective decisions. While there exist …

  83. arXiv cs.LG TIER_1 English(EN) · Mintae Kim, Koushil Sreenath ·

    In-Distribution Imagination for Model-Based Offline Reinforcement Learning

    arXiv:2609.38673v1 Announce Type: new Abstract: Model-based offline reinforcement learning (MBORL) improves sample efficiency through model-generated trajectories. However, accumulative model error can drive imagined trajectories outside the offline data distribution, leading to …

  84. arXiv cs.LG TIER_1 English(EN) · Xinyi Ni, Lifeng Lai ·

    Robust Risk-Sensitive Reinforcement Learning from Corrupted Human Feedback

    arXiv:2609.38938v1 Announce Type: new Abstract: Reinforcement learning with human feedback (RLHF) learns from human comparisons, which can be corrupted or deliberately manipulated. This paper studies online risk-sensitive RLHF with static conditional value-at-risk (CVaR) under ad…

  85. arXiv cs.LG TIER_1 English(EN) · Kihyun Yu, Seoungbin Bae, Dabeen Lee ·

    Learning Infinite-Horizon Average-Reward CMDPs via State Augmentation

    arXiv:2609.39093v1 Announce Type: new Abstract: We study infinite-horizon average-reward constrained Markov decision processes (CMDPs) under the weakly communicating assumption. Existing high-probability guarantees for this setting either require computationally inefficient algor…

  86. arXiv cs.LG TIER_1 English(EN) · Shuxiao Xie, Shuyang Xie, Dezhi Ran, Wei Yang, Tao Xie ·

    Backward-State Policy Is Part of the Learning Algorithm

    arXiv:2609.39813v1 Announce Type: new Abstract: Low-precision training rounds tensors that the backward pass reads again, often for several gradients; each use can read the forward's rounded value, the original, or a new random rounding. This backward-state policy looks like a me…

  87. arXiv cs.LG TIER_1 English(EN) · Seonvin Cho, Soohyun Choi, Songnam Hong ·

    Role-Adaptive Policy Optimization for Offline Reinforcement Learning

    arXiv:2609.40149v1 Announce Type: new Abstract: Policy regularization in offline reinforcement learning balances policy improvement against reliance on uncertain value estimates. This balance can differ between selecting actions for execution and supplying actions for critic boot…

  88. arXiv cs.LG TIER_1 English(EN) · Buse Y{\i}lmaz ·

    Reinforcement Learning-Guided Graph Transformations for SpTRSV Optimization

    arXiv:2609.40159v1 Announce Type: cross Abstract: Sparse triangular solve (SpTRSV) is a fundamental kernel in numerous scientific and engineering applications. However, the data dependencies inherent in sparse triangular matrices significantly limit the available parallelism and …

  89. arXiv cs.LG TIER_1 English(EN) · Shuai Dong, Yongfu Zhu, Yuqi Xu, Weichu Xie, Liuwenpu, Ziyue Wang, Kaiwen Tuo, Congcong Wang, Siyuan Wang, Wenqi Shao, Shuai Yang, Ji Zhao, Caoyuan Ma, Wenzheng Chang, Taiqiang Wu, Xinlei Yu, Hongrui Wu, Xiaoxuan He, Fangke Chen, Dianyi Wang, Kanghui Ti… ·

    RL Starts before RL: On Policy Distillation for Better Reinforcement Learning

    arXiv:2609.28145v2 Announce Type: replace Abstract: Reinforcement learning (RL) improves reasoning, but its performance depends on the policy from which training begins. We study on-policy distillation (OPD) as a preparation stage for RL and ask whether its benefits extend beyond…

  90. Hugging Face Daily Papers TIER_1 English(EN) ·

    MetaRubric: Learning to Reward for Rubric-Based Reinforcement Learning

    Rubric-based reinforcement learning extends reward-driven optimization to open-ended tasks by assigning partial credit to individual response requirements. However, rubric judges can assign a high criterion score even when the information or action it requires is absent from the …

  91. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Qilei Li ·

    After Cooperation Is Learned: Gradient Routing and Optimizer-Dependent Maintenance in Multi-Agent Reinforcement Learning

    Cooperative MARL is commonly evaluated through cooperation discovery from random initialization, leaving open whether continued optimization can destabilize learned cooperation. Actor-critic comparisons can also conflate critic presence with value gradients entering shared actor …

  92. arXiv cs.NE (Neural & Evolutionary) TIER_1 English(EN) · Sebastian Risi ·

    Continual Reinforcement Learning with Neuroevolution

    Despite many studies about causes and remedies of plasticity loss in Reinforcement Learning (RL) under continual task changes, no RL method has yet consistently achieved a good balance between adaptation and forgetting. Here we turn to an alternative optimization paradigm, neuroe…

  93. arXiv cs.NE (Neural & Evolutionary) TIER_1 English(EN) · Sebastian Risi ·

    Continual Reinforcement Learning with Neuroevolution

    Despite many studies about causes and remedies of plasticity loss in Reinforcement Learning (RL) under continual task changes, no RL method has yet consistently achieved a good balance between adaptation and forgetting. Here we turn to an alternative optimization paradigm, neuroe…

  94. arXiv cs.AI TIER_1 English(EN) · Yang chen, Yitan Zhang, Michael Witbrock, Shuyue Hu ·

    Loop-Free Inverse Reinforcement Learning via Sequential Value Recovery with Q-Score Matching

    arXiv:2609.38955v1 Announce Type: cross Abstract: Inverse Reinforcement Learning (IRL) aims to recover a reward function that explains expert demonstrations. Existing IRL methods typically rely on a bi-level optimization procedure that alternates between reward learning and polic…

  95. arXiv cs.AI TIER_1 English(EN) · Yitong Qiao, Tiantian He, Lei Liu, Yue Shen, Jian Wang, Jinjie Gu, Zhixuan Chu ·

    From Imitation to Reward Discovery: On-Policy Warmup for Agentic RL

    arXiv:2609.39436v1 Announce Type: cross Abstract: Reinforcement learning with a verifiable reward (RLVR) offers a scalable approach to training language-model agents, yet sparse outcome rewards can leave early training with little signal for policy improvement. We identify an On-…

  96. arXiv cs.CL TIER_1 English(EN) · Junyu Lu, Shichao Weng, Zhiqiang Wang, Haojie Luo, Jingfan Zhang, Yuhua Zhou, Cheng Du, Yuzhuo Zhang, Xi Li, Jinwei Du, Tiancheng Feng, Chuan Xiao, Shuyuan Zheng ·

    OPTS-TTPO: Enhancing Finite-Sample Policy-Gradient Learning with Tree Search

    arXiv:2609.40035v1 Announce Type: new Abstract: The policy-gradient theorem gives the exact gradient under the current policy, but finite on-policy samples may miss rare high-return trajectories. We study whether tree search improves their coverage within a fixed budget while con…

  97. arXiv cs.AI TIER_1 English(EN) · Yifan Zhang, Qifan Zhang, Liang Zheng ·

    Completion-Aware Cross-Fidelity Offline-to-Online Reinforcement Learning for Multi-Line Bus Holding

    arXiv:2609.39868v1 Announce Type: new Abstract: Exploratory reinforcement learning (RL) on an operating bus fleet is impractical,while policies trained only from historical data cannot acquire new experience. Hybrid Offline-and-Online (H2O) RL combines fixed target replay with si…

  98. arXiv cs.AI TIER_1 English(EN) · Ali Aouad, Aymane El Gadarri, Vivek F. Farias ·

    Revisiting scaling laws for reward optimization

    arXiv:2609.38526v1 Announce Type: cross Abstract: Scaling laws for optimization against reward models in AI alignment have pinned down how performance depends on optimization effort---measured by a KL-divergence budget relative to a reference policy. Beyond a certain budget, over…

  99. arXiv cs.AI TIER_1 English(EN) · Seokmin Ko, Taewon Goo, Kihyuk Hong ·

    DAMPER: Return-Prioritized Gradient Control for Smooth Policies

    arXiv:2609.38903v1 Announce Type: cross Abstract: Actor-critic methods achieve strong performance in continuous control, but their policies can produce highly oscillatory actions. A common remedy is to add auxiliary smoothness losses. However, their contribution can be negligible…

  100. arXiv cs.AI TIER_1 English(EN) · Yilie Huang, Wenpin Tang, Xunyu Zhou ·

    ART for Diffusion Sampling: A Reinforcement Learning Approach to Timestep Schedule

    arXiv:2601.18681v3 Announce Type: replace-cross Abstract: We consider time discretization for score-based diffusion models to generate samples from a learned reverse-time dynamic on a finite grid. Uniform and hand-crafted grids can be suboptimal given a budget on the number of ti…

  101. arXiv cs.AI TIER_1 English(EN) · Tianle Zhong, Neiwen Ling, Yifan Pi, Zijun Wei, Tianshu Yu, Geoffrey Fox, Peng Wu, Xiao Yu ·

    Diagnosing Training Inference Mismatch in LLM Reinforcement Learning via a Zero-Mismatch Reference

    arXiv:2605.14220v2 Announce Type: replace-cross Abstract: Modern LLM RL systems separate rollout generation from policy optimization. These two stages are expected to produce token probabilities that match exactly. However, implementation differences can make them assign differen…

  102. Hugging Face Daily Papers TIER_1 English(EN) ·

    Reinforcement Learning-Guided Graph Transformations for SpTRSV Optimization

    Sparse triangular solve (SpTRSV) is a fundamental kernel in numerous scientific and engineering applications. However, the data dependencies inherent in sparse triangular matrices significantly limit the available parallelism and make efficient workload distribution challenging. …

  103. Hugging Face Daily Papers TIER_1 English(EN) ·

    Learning Infinite-Horizon Average-Reward CMDPs via State Augmentation

    We study infinite-horizon average-reward constrained Markov decision processes (CMDPs) under the weakly communicating assumption. Existing high-probability guarantees for this setting either require computationally inefficient algorithms or have suboptimal dependence on the numbe…

  104. arXiv cs.LG TIER_1 English(EN) · Sanghyun Hahn, Jonghyun Choi ·

    Action Chunking Proximal Policy Optimization with Feedback Correction

    arXiv:2609.36250v1 Announce Type: new Abstract: Action chunking provides temporal abstraction in reinforcement learning by selecting short action sequences instead of individual actions, but many existing approaches face two limitations in high-dimensional robotic control. First,…

  105. arXiv cs.LG TIER_1 English(EN) · Shengchao Hu, Peng Wang, Jifeng Hu, Qiyang Zhou, Anning Hu, Li Shen, Ya Zhang, Dacheng Tao ·

    Q-learning Penalized Transformer for Safe Offline Reinforcement Learning

    arXiv:2609.34426v2 Announce Type: replace Abstract: This paper addresses the problem of safe offline reinforcement learning, which involves training a policy to satisfy safety constraints using an offline dataset. This problem is inherently challenging as it requires balancing th…

  106. arXiv cs.LG TIER_1 English(EN) · Nida Zamir, I-Hong Hou ·

    Restless Bandits with Individual Penalty Constraints: Near-Optimal Indices and Deep Reinforcement Learning

    arXiv:2604.04101v4 Announce Type: replace Abstract: This paper investigates the Restless Multi-Armed Bandit (RMAB) framework under individual penalty constraints to address resource allocation challenges in dynamic wireless networked environments. Unlike conventional RMAB models,…

  107. arXiv cs.LG TIER_1 English(EN) · Rayan Mazouz, Haibo Zhao, Chris Hillar, Christian Shewmake ·

    Sample Complexity of Equivariant Reinforcement Learning

    arXiv:2609.36421v1 Announce Type: cross Abstract: Reinforcement learning (RL) is a powerful framework for robotic control, yet its practical application is often hindered by high sample complexity. This is particularly restrictive in physical domains where interaction data is cos…

  108. arXiv cs.LG TIER_1 English(EN) · T. Konstantin Rusch, Tim Seyde, Jared Boyer, Zach J. Patterson, Daniela Rus ·

    Looped Actor: Depth-Recurrent Reasoning Models for Reinforcement Learning

    arXiv:2609.37432v1 Announce Type: new Abstract: Looped reasoning models repeatedly apply a shared set of parameters, enabling more computation without increasing the model size. These models also support input-dependent computation by dynamically deciding when to stop looping. Mo…

  109. arXiv cs.LG TIER_1 English(EN) · Zhiwei Wang, Yanxi Chen, Yaliang Li, Bolin Ding ·

    Fine-Tuning on Self-Generated and Reward-Weighted Data: Learning Dynamics, Convergence Rates, and Benefits of Off-Policyness

    arXiv:2609.36945v1 Announce Type: new Abstract: We study the learning dynamics of fine-tuning a policy model on self-generated and reward-weighted data, with particular focus on a generalized version of REINFORCE -- referred to as RE(S) -- that updates the rollout distribution on…

  110. arXiv cs.LG TIER_1 English(EN) · Fanchao Chen, Hengyu Fu, Shivaram Venkataraman, Jiantao Jiao ·

    Where the Model Changes Its Mind: Hindsight-Divergence Localization for Efficient Reinforcement Learning with Verifiable Rewards

    arXiv:2609.36864v1 Announce Type: new Abstract: Group-relative methods for reinforcement learning with verifiable rewards (RLVR) learn from differences in rollout outcomes. Independently sampling complete trajectories is costly and does not explicitly explore the decision space a…

  111. arXiv cs.LG TIER_1 English(EN) · Zhaojun Peng ·

    Markovian Nonconvex ADMM for Reinforcement Learning: Bellman-Resolvent Stability Beyond Smooth Blocks

    arXiv:2609.36859v1 Announce Type: new Abstract: We identify and study a structural mechanism for Markovian nonconvex ADMM in reinforcement learning. Using finite discounted MDPs as a canonical proving ground, we show that the discounted Bellman resolvent $(I-\gamma P_\pi)^{-1}$ c…

  112. arXiv cs.LG TIER_1 English(EN) · Ruichuan Huang, Jinghan Liu, Congliang Chen ·

    Towards Better Training Signal: Advantage Clipped Policy Optimization

    arXiv:2609.36816v1 Announce Type: new Abstract: Reinforcement learning (RL) has become a cornerstone for improving the reasoning capabilities of large language models (LLMs), but the need for on-policy data substantially limits training efficiency. Reusing off-policy data through…

  113. arXiv cs.LG TIER_1 English(EN) · Zijun Chen, Zihan Zhang ·

    Optimal Multi-Reward Reinforcement Learning

    arXiv:2609.36486v1 Announce Type: new Abstract: We study an unknown-transition finite-horizon Markov decision process (MDP) with a finite collection of known reward functions $\{r^1, r^2, \ldots, r^M\}$. The goal is to output an $\epsilon$-optimal policy for every reward using on…

  114. arXiv cs.LG TIER_1 English(EN) · Hsiao-Ru Pan, Florent Draye, Bernhard Sch\"olkopf ·

    ABC: Advantage-Based Control Variates for Reinforcement Learning with Verifiable Rewards

    arXiv:2609.36058v1 Announce Type: new Abstract: Recent progress in reinforcement learning with verifiable rewards (RLVR) has highlighted the effectiveness of simple critic-free policy-gradient methods such as Group Relative Policy Optimization (GRPO). In contrast, actor-critic me…

  115. arXiv cs.LG TIER_1 English(EN) · Tianwei Ni, Vineet Jain, Akash Karthikeyan, Pierre-Luc Bacon ·

    From Static Policies to Adaptive Priors in Offline Reinforcement Learning

    arXiv:2609.35880v1 Announce Type: new Abstract: Offline reinforcement learning (RL) has traditionally focused on learning policies for direct deployment under conservative objectives, where uncertainty outside the offline dataset is treated pessimistically to ensure robustness. W…

  116. arXiv cs.CL TIER_1 English(EN) · Selim Furkan Tekin, Gaowen Liu, Ramana Rao Kompella, Ling Liu ·

    Dynamic Optimizations of LLM Ensembles with Two-Stage Reinforcement Learning Agents

    arXiv:2502.04492v3 Announce Type: replace Abstract: The advancement of LLMs and their accessibility have triggered renewed interest in multi-agent reinforcement learning as robust and adaptive frameworks for dynamically changing environments. This paper introduces \texttt{RL-Foca…

  117. arXiv cs.AI TIER_1 English(EN) · Emiliano Penaloza, Dane Malenfant, Dheeraj Vattikonda, Roger Creus Castanyer, Siddarth Venkatraman, Abhay Puri, Jonathan Light, Matthew James Sargent, Augustine N. Mavor-Parker, Massimo Caccia, Lucas Caccia, Glen Berseth, Esmeralda S. Whitammer, Alessand… ·

    KV-streams for Efficient Compaction in Agentic Reinforcement Learning

    arXiv:2609.35750v2 Announce Type: replace-cross Abstract: Scaling the horizon of agentic LLMs is bottlenecked by the need to fit ever longer context traces in GPU memory. Context compaction has been the most popular mechanism to alleviate this issue, keeping GPU memory constant f…

  118. arXiv cs.AI TIER_1 English(EN) · Minung Kim, Jeongmo Kim, Gwanwoo Choi, Seungyul Han ·

    Generative Support Realignment for Cross-Domain Offline Reinforcement Learning

    arXiv:2605.13054v2 Announce Type: replace-cross Abstract: Cross-domain offline reinforcement learning learns a target policy from pre-collected source and target datasets with different dynamics. When target data are scarce, effectively compensating for their limited coverage usi…

  119. arXiv cs.AI TIER_1 English(EN) · Xian Yu, Minheng Xiao, Lei Ying ·

    On the Approximation and Convergence of Distributional Policy Gradient Algorithms for Risk-Sensitive Reinforcement Learning

    arXiv:2405.14749v3 Announce Type: replace-cross Abstract: Risk-sensitive reinforcement learning (RL) is crucial for maintaining reliable performance in high-stakes applications. While traditional RL methods aim to learn a point estimate of the random cumulative cost, distribution…

  120. arXiv cs.AI TIER_1 English(EN) · Xudong Chen, Yixin Liu, Hua Wei, Kaize Ding ·

    COLLATOR: Compositional Multi-Agent Orchestration with Counterfactual Reinforcement Learning

    arXiv:2605.14483v2 Announce Type: replace Abstract: Large language models (LLMs) provide a flexible foundation for multi-agent systems, but their effectiveness and computational cost depend critically on orchestration design. Across different tasks, role design, capacity assignme…

  121. arXiv cs.AI TIER_1 English(EN) · Hongyang Li, Xiao Li, Caesar Wu, Said Mammar, Gr\'egoire Danoy, Pascal Bouvry ·

    Unlocking the Critic: Reward-Free Policy Optimization for LLM Post-Training

    arXiv:2609.37119v1 Announce Type: cross Abstract: Recent approaches to reinforcement learning (RL) post-training for large language models increasingly remove the critic to reduce training instability and memory overhead. Even where a critic is trained, it is discarded once train…

  122. arXiv cs.AI TIER_1 English(EN) · Zihao Chen, Fanxiang Xiong, Hongran Ren, Xuefeng Bai, Zhongxiang Dai, Kehai Chen, Zhiguo Zhang, Zhiyong Wang, Yu Cheng ·

    SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning

    arXiv:2609.36552v1 Announce Type: cross Abstract: Maximum Likelihood Reinforcement Learning (MaxRL) targets prompt-wise log-success and has shown strong performance on reasoning tasks. Under finite rollout budgets, however, the estimator used by MaxRL attenuates each prompt's lik…

  123. arXiv cs.AI TIER_1 English(EN) · Muhang Tian, Sherry Yang ·

    Reward-rate Policy Gradient for Efficient Machine Learning Engineering Agents

    arXiv:2609.36393v1 Announce Type: cross Abstract: Traditional reinforcement learning (RL) techniques focus on maximizing expected cumulative reward, where each action assumes to take a constant unit of time. However, this assumption does not hold for agentic RL tasks such as mach…

  124. arXiv cs.AI TIER_1 English(EN) · Muhammad U. Nasir, Alex Vogt, Steven D. James, Julian Togelius ·

    State Trace Rationale As Auxiliary Task in Reinforcement Learning

    arXiv:2609.36867v1 Announce Type: new Abstract: We propose STRAT, an auxiliary task that trains deep reinforcement learning (RL) agents to predict a short textual trace of their own state. Inspired by human spatial navigation, the description combines landmark, route, and survey …

  125. arXiv cs.AI TIER_1 English(EN) · Zhenrui Yue, Huimin Zeng, Yueqi Wang, Yaokun Liu, Fengran Mo, Jinghan Zhang, Mung Yao Jia, Gyuseok Lee, Yang Zhang, Na Wei, Dong Wang ·

    SIPO: Unifying Reinforcement Learning with On-Policy Self-Distillation

    arXiv:2609.36742v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for improving large language models (LLMs) on various tasks, yet its sparse outcome rewards lack token-level credit assignment for intermediate ste…

  126. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Tuomas Sandholm ·

    Regularized policy gradient with learned mixtures of Gaussians for games with continuous actions

    Most successes of superhuman game-playing algorithms are in games with discrete actions, yet in auctions, robotics, sports, or trading, actions are nearly continuous. Prior techniques either rely on expert-designed discretizations or are sample inefficient. We present a scalable …

  127. arXiv cs.LG TIER_1 English(EN) · Yichen Lin, Xuyuan Xiong, Xue Wang, Xiangfu Meng, Mike Mingcheng Wei, Tao Yao ·

    From Weak Data to Strong Policy: Q-Targets Enable Provable In-Context Reinforcement Learning

    arXiv:2609.30391v1 Announce Type: new Abstract: Existing in-context reinforcement learning methods mainly pretrain Transformers with supervised behavior-prediction objectives. This enables task inference from context, but makes the learned policy strongly depend on the quality of…

  128. arXiv cs.AI TIER_1 English(EN) · Mattie Terzolo, Mikolaj Sacha, Ayan Sinha, Andrew Rabinovich ·

    Reinforcement Learning with Decomposed Subtasks

    arXiv:2609.27035v2 Announce Type: replace Abstract: Group Relative Policy Optimization (GRPO) and related policy-gradient methods for training language model agents collapse an entire multi-turn rollout into a single scalar trajectory reward before it enters the policy update. Wh…

  129. Hugging Face Daily Papers TIER_1 English(EN) ·

    Where the Model Changes Its Mind: Hindsight-Divergence Localization for Efficient Reinforcement Learning with Verifiable Rewards

    Group-relative methods for reinforcement learning with verifiable rewards (RLVR) learn from differences in rollout outcomes. Independently sampling complete trajectories is costly and does not explicitly explore the decision space at critical positions. Feedback on a completed tr…

  130. Hugging Face Daily Papers TIER_1 English(EN) ·

    Q-learning Penalized Transformer for Safe Offline Reinforcement Learning

    This paper addresses the problem of safe offline reinforcement learning, which involves training a policy to satisfy safety constraints using an offline dataset. This problem is inherently challenging as it requires balancing three highly interconnected and competing objectives: …

  131. Hugging Face Daily Papers TIER_1 English(EN) ·

    Agent Priors-guided Policy Learning

    Robots that learn from a few demonstrations often require two forms of generalization. Compositional generalization recombines skills to solve new tasks, and skill generalization lets the learned policy behind each skill work in new situations. The two depend on each other, yet i…

  132. Hugging Face Daily Papers TIER_1 English(EN) ·

    The Low-Rank Structure of VLA Reinforcement Learning

    Reinforcement learning (RL) is increasingly used to post-train vision-language-action (VLA) models, yet how RL reshapes these policies remains poorly understood. We find that RL across widely used flow-based VLA models, including π_{0.5} and GR00T~N1.5/N1.6, on LIBERO, ManiSkill,…

  133. Hugging Face Daily Papers TIER_1 English(EN) ·

    Label-free steering: Compressing test-time reinforcement learning into bias-only subspaces

    Test-time reinforcement learning (TTRL) enables models to improve their reasoning without relying on labeled training data, but existing approaches typically optimize a large fraction of the model parameters. This raises a natural question: can effective test-time adaptation emer…

  134. Hugging Face Daily Papers TIER_1 English(EN) ·

    Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning

    Unified multimodal models can both look at and render images, so in principle they can repair their own generations: diagnose what an image gets wrong, revise it, observe the result, and diagnose again. Whether a revision helps is known only after it is rendered, so the reflectio…

  135. Hugging Face Daily Papers TIER_1 English(EN) ·

    QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents

    Large language model (LLM) agents increasingly undertake extreme-long (xlong) horizon tasks, where a single execution can span hours, hundreds of model--environment interactions, and nearly 1M tokens per rollout. Applying online reinforcement learning (RL) to such executions pose…

  136. Hugging Face Daily Papers TIER_1 English(EN) ·

    Playing to Par: Reinforcement Learning for Provably Optimal Quadrilateral Block Decompositions

    A quadrilateral block decomposition of a planar domain is judged by whether it is complete, whether its elements are well shaped, and how many of its vertices are irregular. The last has a provable floor: the discrete Gauss-Bonnet identity enforces a lower bound on the total vert…

  137. arXiv cs.AI TIER_1 English(EN) · James Wu, Chris R. Sims ·

    Policy Complexity, Reaction Time, and Bounded Rationality in Reinforcement Learning

    arXiv:2609.28737v1 Announce Type: cross Abstract: Biological agents do not learn under conditions of unlimited computation. For humans, learning and choice are shaped by constraints on perception, attention, and working memory, which limit how much state information guides behavi…

  138. arXiv stat.ML TIER_1 English(EN) · Akira Kitaoka ·

    Outperformance Inverse Optimization: Learning Objective Functions that Outperform Agent Decisions

    arXiv:2610.09890v1 Announce Type: cross Abstract: Inverse optimization estimates the weights of an objective function that explain observed decisions as optimal solutions, and is used in a variety of fields. For mixed-integer linear programs (MILPs), existing methods aim to repro…

  139. arXiv cs.CV TIER_1 English(EN) · Bingyang Cui, Yujie Zhang, Yiling Xu, Yunfeng Guan ·

    OuroReward: Sequential Reward Scheduling for Reinforcement Learning in Text-to-3D Generation

    arXiv:2610.03423v1 Announce Type: new Abstract: Reinforcement learning (RL) for Text-to-3D (T23D) generation requires optimization across multiple quality dimensions such as semantic alignment and texture clarity. Existing methods typically optimize these dimensions simultaneousl…

  140. arXiv stat.ML TIER_1 English(EN) · Allen Tran, Jia Wan, Nathan Kallus, Aur\'elien Bibaut ·

    Scalable Multi-Task Inverse Reinforcement Learning

    arXiv:2610.00758v1 Announce Type: cross Abstract: By learning transferable rewards, inverse reinforcement learning (IRL) enables counterfactual evaluation of agents under modified environments. Such transfer places strict requirements on coverage since target environments affect …

  141. arXiv cs.CV TIER_1 English(EN) · Yifan Wang, Gordon Guocheng Qian, Yanyu Li, Anil Kag, Yun Fu ·

    Token-Level Video Reinforcement Learning

    arXiv:2610.01973v1 Announce Type: new Abstract: Reinforcement learning (RL) for video generation usually assigns one scalar reward to an entire sampled video. Yet a video is not uniformly flawed: some visual tokens may already satisfy the prompt, whereas others require correction…

  142. arXiv stat.ML TIER_1 English(EN) · Shashank Gupta, Pilhwa Lee ·

    Meta-reinforcement learning with minimum attention

    arXiv:2505.16741v5 Announce Type: replace-cross Abstract: Minimum attention applies the least action principle in changes of control concerning state and time, first proposed by Brockett. The involved regularization is highly relevant in emulating biological control, such as moto…

  143. arXiv stat.ML TIER_1 English(EN) · Toru Kitagawa, Shosei Sakaguchi, Aleksey Tetenov ·

    Constrained Classification and Policy Learning

    arXiv:2106.12886v3 Announce Type: replace-cross Abstract: Modern machine learning approaches to classification, including AdaBoost, support vector machines, and deep neural networks, utilize surrogate loss techniques to circumvent the computational complexity of minimizing empiri…

  144. arXiv stat.ML TIER_1 English(EN) · Deqian Kong, Guangyan Sun, Sheng Cheng, Sirui Xie, Bo Pang, Jianwen Xie, Tony Geng, Caiwen Ding, Ying Nian Wu ·

    Learning to Plan from Random Exploration

    arXiv:2609.38383v1 Announce Type: cross Abstract: Random exploration reveals how an environment can be traversed before a goal is specified. Can this experience support long-range planning without policy-improvement training? Our random-walk analysis explains what temporal relati…

  145. arXiv stat.ML TIER_1 English(EN) · Dylan J. Foster, Alexander Rakhlin ·

    Foundations of Reinforcement Learning and Interactive Decision Making

    arXiv:2312.16730v2 Announce Type: replace-cross Abstract: Interactive decision making is the problem of learning to act well in an unknown environment, using the data that one's own actions generate to continuously improve, and arises in situations ranging from online platforms a…

  146. Towards AI TIER_1 English(EN) · Mayurnayak ·

    Reinforcement Learning from Scratch Part — 1: Introduction

    <figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*bzw_rrScey_iRtRBlLodOg.png" /></figure><p>This is the Part-1 of Series <strong>Reinforcement Learning from Scratch. </strong>In this series we plan to cover the core concepts of modern RL from Q-table, Deep-Q Lea…

  147. Medium — Claude tag TIER_1 English(EN) · Arjun Keerthi ·

    Agentic Finetuning with Reinforcement Learning

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@arjunkeerthi1529.edu/agentic-finetuning-with-reinforcement-learning-b42e68857fbb?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1200/0*lNH-_f93luam3niK" width="1200" /…

  148. dev.to — LLM tag TIER_1 English(EN) · Prabhakar Chaudhary ·

    LoGRA: How Low-Rank Gradient Sketches Make LLM Reinforcement Learning Fit on Real Hardware

    <h1> LoGRA: How Low-Rank Gradient Sketches Make LLM Reinforcement Learning Fit on Real Hardware </h1> <p>Reinforcement learning post-training has become one of the most effective ways to improve large language model reasoning. Models like DeepSeek-R1 demonstrated that RL-based fi…

  149. dev.to — LLM tag TIER_1 English(EN) · Seth Wheeler ·

    Reading a Reinforcement Learning Taxonomy Against My Own Tools

    <p>Somebody handed me a list of reinforcement learning algorithms and asked whether any of them could help my tools. It is a ChatGPT session printed to PDF: roughly ninety algorithm names across fifteen sections, with no descriptions, no citations and no results. I want to be pre…