English(EN)KV-streams for Efficient Compaction in Agentic Reinforcement Learning
AI 研究在数学、适应性和公平性方面推进强化学习 · 跟踪 10 个来源
作者PulseAugur 编辑部·[149 个来源]·
研究人员正在探索先进的强化学习技术,以提高 AI 的数学推理和适应能力。一篇论文介绍了函数结构图强化学习 (FSG-RL),通过从函数图生成代码并使用多验证器反馈来增强数学问题解决能力,显著提高了准确性。其他研究侧重于通过样本重用来优化近端策略优化 (PPO),并探索用于测试时适应的无数据方法。此外,还在开发无监督强化学习、具有神经进化的持续学习以及将公平性与 AI 系统的可解释性相结合的新方法。
AI
arXiv:2407.18994v2 Announce Type: replace-cross Abstract: We consider the automatic online synthesis of black-box test cases from functional requirements specified as automata for reactive implementations. The goal of the tester is to reach some given state, so as to satisfy a co…
arXiv cs.CL
TIER_1English(EN)·Zirui Li, Rech Silas, Lauri Juvela, Tom Backstrom, Mikko Kurimo·
arXiv:2609.27599v2 Announce Type: cross Abstract: Generating controllable and human-like emphasis remains an open challenge in text-to-speech, even when explicit emphasis control signals are provided in the text input, limiting the communicative accuracy of synthetic speech in re…
arXiv:2610.08963v1 Announce Type: cross Abstract: Asynchronous reinforcement learning (RL) for large language model (LLM) agents trains one policy on trajectories generated by another: rollouts come from stale checkpoints, and the inference engine's probabilities differ from the …
arXiv cs.CL
TIER_1English(EN)·Shaokun Zhang, Yifan Zhang, Jian Hu, Yueying Li, Hao Zhang, Binfeng Xu, Jan Kautz, Yi Dong·
arXiv:2610.06647v2 Announce Type: replace Abstract: Reinforcement learning has greatly advanced the capabilities of large language models, but its memory demands remain a barrier to broader adoption. We introduce LoGRA, an approach to RL post-training that reduces memory by retai…
arXiv cs.LG
TIER_1English(EN)·Shripad V. Deshmukh, Yaswanth Chittepu, Dhawal Gupta, Philip Thomas, Scott Niekum·
arXiv:2610.09108v1 Announce Type: new Abstract: Policy learning drives many of the most consequential and heavily-invested applications of reinforcement learning today. Yet the core optimization problem it rests on (maximizing expected return) is notoriously non-convex, even unde…
arXiv cs.LG
TIER_1English(EN)·Michael Tang, Mahmoud Abdelgalil, Jorge I. Poveda·
arXiv:2610.09120v1 Announce Type: new Abstract: Randomized exploration is central to bandit learning, multi-agent reinforcement learning, and zeroth-order policy search, yet its independence and privacy are usually only treated as technical assumptions. We show that these propert…
arXiv cs.LG
TIER_1English(EN)·John L. Zhou, Yuxuan Dong, Jonathan C. Kao·
arXiv:2610.09247v1 Announce Type: new Abstract: The difficulty of learning goal-reaching policies is often attributed to a "curse of horizon" that manifests as bias accumulation in temporal-difference backups and noisy advantage estimates. In this work, we identify an additional …
arXiv:2610.09597v1 Announce Type: new Abstract: Asynchronous RL accelerates large language model post-training by decoupling rollout generation from optimization, but trains on stale trajectories. Existing methods primarily correct token-level policy mismatch through importance-r…
arXiv cs.LG
TIER_1English(EN)·Mahmoud Selim, Cristina Cipriani, Karl Henrik Johansson·
arXiv:2610.09763v1 Announce Type: new Abstract: Offline reinforcement learning enables reward-driven policy improvement from fixed datasets without requiring online exploration, making it particularly attractive in safety-critical domains. A central challenge, however, is distrib…
arXiv:2610.09914v1 Announce Type: new Abstract: Reinforcement learning is crucial for improving large language models' reasoning and generalization. It relies on massive rollouts whose lengths become increasingly long-tailed as context windows grow. In on-policy training, these l…
arXiv cs.LG
TIER_1English(EN)·Tommaso Marzi, Ahmed Hendawy, Jan Peters, Carlo D'Eramo, Andrea Cini, Cesare Alippi·
arXiv:2610.10302v1 Announce Type: new Abstract: In Continual Multi-Agent Reinforcement Learning (CMARL), agents learn cooperative policies across sequences of tasks, aiming to adapt effectively to new tasks while preserving the ability to solve previously encountered ones. In man…
arXiv:2610.10326v1 Announce Type: new Abstract: We study learning optimal policies in average-reward multichain Markov decision processes (MDPs), where the optimal gain may depend on the initial state and recurrence structures vary across policies, creating challenges for reinfor…
arXiv:2610.09337v1 Announce Type: cross Abstract: Policy-based reinforcement learning (RL) approaches have produced promising results for autonomous cyber defense; however, they are sample-inefficient in settings where defenders must respond under delayed, partial observations wi…
arXiv:2610.09712v1 Announce Type: cross Abstract: We develop a continuous-time deterministic policy gradient framework for reinforcement learning with reflected state dynamics, where the state process is governed by a controlled reflected stochastic differential equation on a hyp…
arXiv:2610.10167v1 Announce Type: cross Abstract: Policy learning in digital experimentation faces three challenges: weak signal-to-noise ratios, rich covariate spaces, and massive data volumes. We formalize this regime by modeling treatment-effect estimates from increasingly fin…
arXiv:2602.08234v2 Announce Type: replace Abstract: Large Language Model (LLM) agents have shown stunning results in complex tasks, yet they often operate in isolation, failing to learn from past experiences. Existing memory-based methods primarily store raw trajectories, which a…
arXiv:2610.06883v1 Announce Type: cross Abstract: Neural operators provide fast surrogates for time-dependent PDEs, but autoregressive deployment creates a refinement-allocation problem: prediction errors vary over space and time, while only a finite number of local corrections c…
arXiv cs.LG
TIER_1English(EN)·Silong Yong, Stephen Sheng, Carl Qi, Xiaojie Wang, Evan Sheehan, Anurag Shivaprasad, Yaqi Xie, Katia Sycara, Yesh Dattatreya·
arXiv:2604.00055v2 Announce Type: replace-cross Abstract: Existing robotic foundation policies are trained primarily via large-scale imitation learning. While such models demonstrate strong capabilities, they often struggle with long-horizon tasks due to distribution shift and er…
arXiv cs.LG
TIER_1English(EN)·Valliappan Chidambaram Adaikkappan, Sai Rajeswar, Pietro Mazzaglia, David Meger·
arXiv:2605.09364v2 Announce Type: replace Abstract: This paper investigates robust representation learning in offline goal-conditioned reinforcement learning (GCRL). Particularly in sparse reward scenarios, learning representations that align state and goal latents is a challenge…
arXiv:2610.07814v1 Announce Type: cross Abstract: How far can stochastic gradient descent ascent (SGDA) go by tuning its timescale ratio and step sizes in nonconvex min-max games? We answer this question for nonconvex-PL (NC-PL) games by establishing the first tight complexity of…
arXiv cs.LG
TIER_1English(EN)·Fernando Martinez, Tao Li, Yingdong Lu, Juntao Chen·
arXiv:2610.07704v1 Announce Type: cross Abstract: Fully decentralized multi-agent reinforcement learning (MARL), also referred to as independent learning, requires each agent to learn and act using only its local information and experience, without a centralized critic or inter-a…
arXiv:2610.08561v1 Announce Type: new Abstract: Reinforcement learning (RL) has become a standard tool for post-training language models on reasoning tasks, where the policy is updated by reward feedback while exploring the space of responses. Despite its empirical success, theor…
arXiv:2610.08271v1 Announce Type: new Abstract: Classical reinforcement learning (RL) assumes that a reward is observed for every visited state-action pair. However, in real-world applications such as autonomous driving, such fine-grained feedback can be costly or difficult to co…
arXiv:2610.07910v1 Announce Type: new Abstract: Deep Reinforcement Learning policies can produce nonsmooth action oscillations that hinder deployment on physical robots. Existing architectural and penalty-based approaches seek spatial smoothness by directly reducing sensitivity t…
arXiv:2610.07332v1 Announce Type: cross Abstract: Long-horizon LLM agents are frequently implemented using sparse mixture-of-experts (MoE) models, yet the co-design of agentic behavior and MoE structures remains underexplored. In this work, we comprehensively study the connection…
arXiv cs.AI
TIER_1English(EN)·Chenliang Li, Adel Elmahdy, Alex Boyd, Zhongruo Wang, Siliang Zeng, Alfredo Garcia, Parminder Bhatia, Taha Kass-Hout, Cao Xiao, Mingyi Hong·
arXiv:2511.20718v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) algorithms such as PPO and GRPO are widely used to train large language models (LLMs) for multi-turn agentic tasks. However, in off-policy training pipelines, these methods can exhibit unstable …
arXiv:2610.08743v1 Announce Type: cross Abstract: Sequential recommenders typically use a fixed slate size even though the number of useful alternatives changes within a session. We propose Reinforcement Learning with Calibrated Pruning (RLCP), which adapts the retained action se…
arXiv:2610.07899v1 Announce Type: cross Abstract: Generative actors are transforming offline reinforcement learning (RL) by enabling expressive policy classes that model complex action distributions. However, this expressiveness also exposes a key challenge in heterogeneous datas…
arXiv:2610.07731v1 Announce Type: cross Abstract: Dense retrieval models are typically trained with contrastive objectives that learn effective representations but do not directly optimize retrieval metrics or downstream task performance. To address this problem, we introduce REL…
arXiv:2610.07491v1 Announce Type: cross Abstract: When multiple agents share a cost budget, a common Lagrange multiplier can enforce the aggregate constraint but does not determine how its penalty should be allocated across agents. Uniform penalties ignore heterogeneity in the re…
arXiv cs.AI
TIER_1English(EN)·T. Y. Tsui, Zihao Ye, Pengxiang Cai, Yanchao Li, Yuqiang Li, Zhehong Ai·
arXiv:2610.07226v1 Announce Type: cross Abstract: ``What are the irreducible conditions that are sufficient to produce an outcome?'' is one of the most common questions that recur across computation and science. Its answers, the minimal sufficient witnesses, are what we mean by e…
arXiv:2610.07043v1 Announce Type: cross Abstract: Low-precision execution can substantially accelerate reinforcement learning (RL) for large language models, but discrepancies between learner and sampler execution can destabilize policy optimization. In this paper, we characteriz…
Deep Reinforcement Learning policies can produce nonsmooth action oscillations that hinder deployment on physical robots. Existing architectural and penalty-based approaches seek spatial smoothness by directly reducing sensitivity to changes in state inputs, but their broad const…
Dense retrieval models are typically trained with contrastive objectives that learn effective representations but do not directly optimize retrieval metrics or downstream task performance. To address this problem, we introduce RELER (REinforcement LEarning for Retrieval), a reinf…
Fully decentralized multi-agent reinforcement learning (MARL), also referred to as independent learning, requires each agent to learn and act using only its local information and experience, without a centralized critic or inter-agent communication. Such a stringent information s…
Asynchronous reinforcement learning (RL) for large language model (LLM) agents trains one policy on trajectories generated by another: rollouts come from stale checkpoints, and the inference engine's probabilities differ from the trainer's even at identical parameters. Standard r…
arXiv cs.MA (Multiagent)
TIER_1English(EN)·Alexandre M. Bayen·
When multiple agents share a cost budget, a common Lagrange multiplier can enforce the aggregate constraint but does not determine how its penalty should be allocated across agents. Uniform penalties ignore heterogeneity in the rewards agents sacrifice, while agent-specific multi…
Language models are moving beyond generating answers to pursuing long-horizon goals in interactive environments. Post-training these agents requires long, heterogeneous trajectories, and synchronous systems leave learner engines idle until rollout and verification finish. To sque…
arXiv cs.LG
TIER_1Deutsch(DE)·Jinwoo Kim, Shraddha Barke·
arXiv:2610.02846v1 Announce Type: new Abstract: When rewards are sparse, reinforcement learning with verifiable rewards (RLVR) often uses hints or intermediate guidance to generate more successful rollouts. This enrichment biases policy-gradient updates unless corrected via impor…
arXiv:2610.03223v1 Announce Type: cross Abstract: Long-horizon LLM agents are typically trained with sparse outcome rewards, making trajectory-level objectives too coarse to distinguish the contribution of individual decisions. Step-level credit assignment provides finer-grained …
arXiv cs.LG
TIER_1English(EN)·Ganghun Lee, Minji Kim, Minsu Lee, Byoung-Tak Zhang·
arXiv:2610.02545v1 Announce Type: new Abstract: Reward serves as the primary learning signal in reinforcement learning (RL). However, while reward magnitudes are typically held fixed throughout training, their temporal modulation remains underexplored. In this paper, we propose r…
arXiv:2610.02478v1 Announce Type: new Abstract: Reinforcement learning for large language models typically maximizes expected return, adding up the probabilities of all successful trajectories. However, the classical sum formulation can only report how often the model policy succ…
arXiv cs.AI
TIER_1English(EN)·Mingyu Park, Samyeul Noh, Hyun Myung, Donghwan Lee·
arXiv:2610.02801v1 Announce Type: new Abstract: Model-based reinforcement learning (MBRL) achieves strong sample efficiency by planning within learned latent dynamics, yet its performance degrades substantially under unseen visual distractions such as background variations, light…
arXiv:2610.02824v1 Announce Type: new Abstract: Rubric-based reinforcement learning extends reward-driven optimization to open-ended tasks by assigning partial credit to individual response requirements. However, rubric judges can assign a high criterion score even when the infor…
arXiv cs.AI
TIER_1English(EN)·Xinjie Liu, Ruihan Zhao, Anirban Chaudhuri, Cyrus Neary, Ufuk Topcu, David Fridovich-Keil·
arXiv:2610.02505v1 Announce Type: cross Abstract: Policy gradient methods for on-policy reinforcement learning (RL) can become unstable when expensive, scarce target-domain data yield noisy gradient estimates. We address this challenge by complementing limited high-fidelity (HF) …
arXiv cs.LG
TIER_1English(EN)·Juri Pfammatter, Kaixian Qu, Clemens Schwarke, Victor Klemm, Marco Hutter·
arXiv:2610.03395v1 Announce Type: new Abstract: Long-horizon tasks with sparse rewards pose an exploration bottleneck for goal-conditioned reinforcement learning: a policy started from the initial state rarely reaches the goal and receives no learning signal. Reference motions, h…
arXiv:2610.02847v1 Announce Type: cross Abstract: Open-team multi-agent reinforcement learning studies cooperative systems in which agents may join, leave, or be replaced during an episode. In such settings, the team return changes both because agents choose useful actions and be…
arXiv cs.LG
TIER_1English(EN)·Matthew Brun, Xu Andy Sun·
arXiv:2610.03642v1 Announce Type: new Abstract: Success conditioning is a strategy for improving decision-making policies in stochastic environments; it updates a policy by increasing the probability of taking actions that yield successful outcomes. Success conditioning is common…
arXiv:2610.03620v1 Announce Type: new Abstract: Online reinforcement learning (RL) enables robot policies to improve through physical interaction, but the assistance they require changes as their competence evolves. Existing intervention strategies based on offline estimates or f…
arXiv:2605.07393v2 Announce Type: replace Abstract: A fundamental challenge in model-based offline reinforcement learning (RL) lies in the trade-off between generalization and robustness against exploitation errors in out-of-distribution (OOD) regions. The key to resolving this t…
arXiv:2610.02740v1 Announce Type: cross Abstract: Reinforcement learning for long-horizon agents relies on purely retrospective training signals: credit is assigned only after observing environmental consequences, leaving the agent's belief at action time invisible to the gradien…
arXiv cs.AI
TIER_1English(EN)·Saptarshi Nath, Inish M. D'Souza, Antonio Carta, Soheil Kolouri, Andrea Soltoggio·
arXiv:2610.03119v1 Announce Type: cross Abstract: In lifelong reinforcement learning, retaining previously learned policies is not sufficient for effective transfer to a new task. Useful knowledge may be distributed across several prior policies, and its relevance may change as t…
arXiv cs.AI
TIER_1English(EN)·Guilhem Loussouarn, Nancy Nayak, Kin K. Leung·
arXiv:2610.03475v1 Announce Type: cross Abstract: Many reinforcement-learning (RL) problems are non-stationary yet structured and can be decomposed into phases, each with its own transition probabilities and reward functions. When the phase sequence is known, the common solution …
arXiv:2610.02527v1 Announce Type: cross Abstract: Learned visual reward models are increasingly used to optimize robot policies, yet a reward model can score an execution that acts on the wrong object as highly as one that completes the task. We show that optimizing such a reward…
Long-horizon LLM agents are frequently implemented using sparse mixture-of-experts (MoE) models, yet the co-design of agentic behavior and MoE structures remains underexplored. In this work, we comprehensively study the connections between agentic post-training and MoE expert sel…
Low-precision execution can substantially accelerate reinforcement learning (RL) for large language models, but discrepancies between learner and sampler execution can destabilize policy optimization. In this paper, we characterize the interaction between mismatch and the policy-…
Reinforcement learning (RL) has greatly advanced the capabilities of large language models (LLMs), but its memory demands remain a barrier to broader adoption. We introduce LoGRA, an approach to RL post-training that reduces memory by retaining useful learning signals in low-rank…
``What are the irreducible conditions that are sufficient to produce an outcome?'' is one of the most common questions that recur across computation and science. Its answers, the minimal sufficient witnesses, are what we mean by explanations, mechanisms and reasons. These problem…
arXiv cs.LG
TIER_1English(EN)·Morgan Byrd, Maks Sorokin, Robert Wright, Sehoon Ha·
arXiv:2610.00729v1 Announce Type: new Abstract: This paper explores a reward-based policy to achieve zero-shot transfer between source and target environments with completely different observation spaces. While humans can demonstrate impressive adaptation capabilities, deep neura…
arXiv cs.LG
TIER_1English(EN)·Eleni Nisioti, Andrea Cossu, Kathrin Korte, Sebastian Risi·
arXiv:2610.01583v1 Announce Type: cross Abstract: Despite many studies about causes and remedies of plasticity loss in Reinforcement Learning (RL) under continual task changes, no RL method has yet consistently achieved a good balance between adaptation and forgetting. Here we tu…
arXiv:2610.02012v1 Announce Type: new Abstract: Reinforcement learning (RL) is a powerful paradigm for training agents, yet its success rests on domain expertise of human engineers who design informative reward signals for every new task. Unsupervised RL aims to reduce this engin…
arXiv:2610.01729v1 Announce Type: new Abstract: Algorithmic mathematical reasoning requires reliable decomposition, computation, and aggregation. Final-answer rewards provide limited guidance on intermediate errors, while successful execution does not guarantee mathematical corre…
arXiv cs.LG
TIER_1English(EN)·Ryunyi Lee, Kangjun Noh, Somin Kim, Heedong Kim, Kyungwoo Song·
arXiv:2610.01548v1 Announce Type: new Abstract: As the use of large language models (LLMs) expands, post-training has become increasingly important for adapting them to downstream tasks. However, obtaining reliable supervision remains costly, especially in domains without referen…
arXiv cs.LG
TIER_1English(EN)·Alessandro Montenegro, Riccardo Venturelli, Marco Mussi, Matteo Papini, Alberto Maria Metelli·
arXiv:2610.01399v1 Announce Type: new Abstract: Among on-policy deep reinforcement learning methods, Proximal Policy Optimization (PPO) has become the de facto standard, due to its consistently strong empirical performance across diverse application domains. However, on-policy me…
arXiv:2610.00903v1 Announce Type: new Abstract: Test-time reinforcement learning (TTRL) adapts language models on unlabeled test problems using supervision derived from their own samples. This makes the checkpoint's \emph{entry state} consequential: a diffuse policy provides nois…
arXiv cs.LG
TIER_1English(EN)·Bente Hinkenhuis, Seyed Sahand Mohammadi Ziabari, Ali Mohammed Mansoor Alsahag·
arXiv:2610.00035v1 Announce Type: new Abstract: Predicting student performance from educational interaction data requires models that are both accurate and sufficiently transparent to support meaningful intervention, while demographic information introduces an additional risk of …
Open-team multi-agent reinforcement learning studies cooperative systems in which agents may join, leave, or be replaced during an episode. In such settings, the team return changes both because agents choose useful actions and because the active population itself changes. Standa…
arXiv cs.AI
TIER_1English(EN)·Jinhwan Sul, Alex Oshin, Evangelos A. Theodorou·
arXiv:2610.00849v1 Announce Type: new Abstract: Existing approaches to offline goal-conditioned reinforcement learning (GCRL) struggle with long-horizon tasks. Discounting shrinks value differences between distant states until they fall below the function approximation error, lea…
arXiv:2610.01207v1 Announce Type: new Abstract: When training large language models with reinforcement learning, terminal rewards provide little guidance about which steps matter. Common methods for assigning step credit overlook that work built on uncorrected mistakes is wasted …
arXiv:2610.01458v1 Announce Type: new Abstract: Verifier-free reinforcement learning with probability-based rewards offers a promising way to train LLMs on general reasoning tasks where external verifiers are unavailable. Yet the reliability of these rewards, especially in long-h…
arXiv cs.AI
TIER_1English(EN)·Jude Waide, Robert Lieck·
arXiv:2610.01513v1 Announce Type: new Abstract: Long-term dependencies remain a major challenge for sequential decision-making in the field of AI: RNNs suffer from vanishing gradients and the limited expressivity of vector-based hidden states, whilst Transformer-based models are …
arXiv cs.AI
TIER_1English(EN)·Abid Mohamed Nadhir, Ahmad Al Hanbali, Beggas Mounir·
arXiv:2610.02074v1 Announce Type: new Abstract: Privacy-preserving machine learning presents significant deployment challenges on the cloud for intelligent systems with confidential data. Fully Homomorphic Encryption (FHE) offers a compelling solution for secure computation, pres…
arXiv:2610.00388v1 Announce Type: cross Abstract: Reinforcement learning enables large language model (LLM) agents to learn multi-step behaviors through interaction with their environments. However, rewards in many interactive tasks reflect only the final outcome, providing limit…
arXiv cs.AI
TIER_1English(EN)·Oleg Shchendrigin, Egor Cherepanov, Aleksandr I. Panov, Alexey K. Kovalev·
arXiv:2610.00592v1 Announce Type: cross Abstract: In partially observable reinforcement learning (RL), a later observation can make stored information obsolete or change what it implies for the next decision. Memory architectures and benchmarks for RL mostly test retention, the a…
arXiv cs.AI
TIER_1English(EN)·Christos Spyridon Koulouris, Carlo Campajola·
arXiv:2610.00619v1 Announce Type: cross Abstract: In this paper, we extend earlier findings of supra-competitive outcomes in optimal-execution games by identifying a learned punitive mechanism that deters deviations and provides behavioural evidence of collusion. We investigate t…
arXiv cs.AI
TIER_1English(EN)·Yujie Zhu, Charles A. Hepburn, Matthew Thorpe, Giovanni Montana·
arXiv:2610.01413v1 Announce Type: cross Abstract: Reinforcement learning (RL) algorithms frequently compare probability distributions, such as state visitation distributions induced by policies and experts, action distributions from learned policies and offline datasets, or trans…
arXiv:2610.02039v1 Announce Type: cross Abstract: Recent years have witnessed the rapid adoption of reinforcement learning (RL) in large language model (LLM) post-training, with substantial gains in mathematical reasoning and code generation. In practical systems, however, policy…
arXiv:2610.02198v1 Announce Type: cross Abstract: Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic. However, critics are typically trained to predict returns, and accurate value pr…
arXiv cs.AI
TIER_1English(EN)·Debosmita Bhaumik, Julian Togelius, Georgios N. Yannakakis, Ahmed Khalifa·
arXiv:2605.13570v2 Announce Type: replace Abstract: Constraint-based game content generators that learn local constraints from existing content, such as Wave Function Collapse (WFC), can generate visually satisfying game levels but face challenges in optimizing global properties,…
arXiv:2605.07579v3 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) for Large Reasoning Models rests on variance reduction, which requires both a reliable baseline and high prompt diversity within each training batch. This is especially…
arXiv:2609.38598v1 Announce Type: new Abstract: Partially observable environments pose a fundamental challenge in deep reinforcement learning, requiring agents to compress temporal information from observations and maintain a memory to make effective decisions. While there exist …
arXiv cs.LG
TIER_1English(EN)·Mintae Kim, Koushil Sreenath·
arXiv:2609.38673v1 Announce Type: new Abstract: Model-based offline reinforcement learning (MBORL) improves sample efficiency through model-generated trajectories. However, accumulative model error can drive imagined trajectories outside the offline data distribution, leading to …
arXiv:2609.38938v1 Announce Type: new Abstract: Reinforcement learning with human feedback (RLHF) learns from human comparisons, which can be corrupted or deliberately manipulated. This paper studies online risk-sensitive RLHF with static conditional value-at-risk (CVaR) under ad…
arXiv:2609.39093v1 Announce Type: new Abstract: We study infinite-horizon average-reward constrained Markov decision processes (CMDPs) under the weakly communicating assumption. Existing high-probability guarantees for this setting either require computationally inefficient algor…
arXiv cs.LG
TIER_1English(EN)·Shuxiao Xie, Shuyang Xie, Dezhi Ran, Wei Yang, Tao Xie·
arXiv:2609.39813v1 Announce Type: new Abstract: Low-precision training rounds tensors that the backward pass reads again, often for several gradients; each use can read the forward's rounded value, the original, or a new random rounding. This backward-state policy looks like a me…
arXiv:2609.40149v1 Announce Type: new Abstract: Policy regularization in offline reinforcement learning balances policy improvement against reliance on uncertain value estimates. This balance can differ between selecting actions for execution and supplying actions for critic boot…
arXiv:2609.40159v1 Announce Type: cross Abstract: Sparse triangular solve (SpTRSV) is a fundamental kernel in numerous scientific and engineering applications. However, the data dependencies inherent in sparse triangular matrices significantly limit the available parallelism and …
arXiv:2609.28145v2 Announce Type: replace Abstract: Reinforcement learning (RL) improves reasoning, but its performance depends on the policy from which training begins. We study on-policy distillation (OPD) as a preparation stage for RL and ask whether its benefits extend beyond…
Rubric-based reinforcement learning extends reward-driven optimization to open-ended tasks by assigning partial credit to individual response requirements. However, rubric judges can assign a high criterion score even when the information or action it requires is absent from the …
Cooperative MARL is commonly evaluated through cooperation discovery from random initialization, leaving open whether continued optimization can destabilize learned cooperation. Actor-critic comparisons can also conflate critic presence with value gradients entering shared actor …
Despite many studies about causes and remedies of plasticity loss in Reinforcement Learning (RL) under continual task changes, no RL method has yet consistently achieved a good balance between adaptation and forgetting. Here we turn to an alternative optimization paradigm, neuroe…
Despite many studies about causes and remedies of plasticity loss in Reinforcement Learning (RL) under continual task changes, no RL method has yet consistently achieved a good balance between adaptation and forgetting. Here we turn to an alternative optimization paradigm, neuroe…
arXiv cs.AI
TIER_1English(EN)·Yang chen, Yitan Zhang, Michael Witbrock, Shuyue Hu·
arXiv:2609.38955v1 Announce Type: cross Abstract: Inverse Reinforcement Learning (IRL) aims to recover a reward function that explains expert demonstrations. Existing IRL methods typically rely on a bi-level optimization procedure that alternates between reward learning and polic…
arXiv:2609.39436v1 Announce Type: cross Abstract: Reinforcement learning with a verifiable reward (RLVR) offers a scalable approach to training language-model agents, yet sparse outcome rewards can leave early training with little signal for policy improvement. We identify an On-…
arXiv cs.CL
TIER_1English(EN)·Junyu Lu, Shichao Weng, Zhiqiang Wang, Haojie Luo, Jingfan Zhang, Yuhua Zhou, Cheng Du, Yuzhuo Zhang, Xi Li, Jinwei Du, Tiancheng Feng, Chuan Xiao, Shuyuan Zheng·
arXiv:2609.40035v1 Announce Type: new Abstract: The policy-gradient theorem gives the exact gradient under the current policy, but finite on-policy samples may miss rare high-return trajectories. We study whether tree search improves their coverage within a fixed budget while con…
arXiv:2609.39868v1 Announce Type: new Abstract: Exploratory reinforcement learning (RL) on an operating bus fleet is impractical,while policies trained only from historical data cannot acquire new experience. Hybrid Offline-and-Online (H2O) RL combines fixed target replay with si…
arXiv cs.AI
TIER_1English(EN)·Ali Aouad, Aymane El Gadarri, Vivek F. Farias·
arXiv:2609.38526v1 Announce Type: cross Abstract: Scaling laws for optimization against reward models in AI alignment have pinned down how performance depends on optimization effort---measured by a KL-divergence budget relative to a reference policy. Beyond a certain budget, over…
arXiv:2609.38903v1 Announce Type: cross Abstract: Actor-critic methods achieve strong performance in continuous control, but their policies can produce highly oscillatory actions. A common remedy is to add auxiliary smoothness losses. However, their contribution can be negligible…
arXiv:2601.18681v3 Announce Type: replace-cross Abstract: We consider time discretization for score-based diffusion models to generate samples from a learned reverse-time dynamic on a finite grid. Uniform and hand-crafted grids can be suboptimal given a budget on the number of ti…
arXiv:2605.14220v2 Announce Type: replace-cross Abstract: Modern LLM RL systems separate rollout generation from policy optimization. These two stages are expected to produce token probabilities that match exactly. However, implementation differences can make them assign differen…
Sparse triangular solve (SpTRSV) is a fundamental kernel in numerous scientific and engineering applications. However, the data dependencies inherent in sparse triangular matrices significantly limit the available parallelism and make efficient workload distribution challenging. …
We study infinite-horizon average-reward constrained Markov decision processes (CMDPs) under the weakly communicating assumption. Existing high-probability guarantees for this setting either require computationally inefficient algorithms or have suboptimal dependence on the numbe…
arXiv:2609.36250v1 Announce Type: new Abstract: Action chunking provides temporal abstraction in reinforcement learning by selecting short action sequences instead of individual actions, but many existing approaches face two limitations in high-dimensional robotic control. First,…
arXiv cs.LG
TIER_1English(EN)·Shengchao Hu, Peng Wang, Jifeng Hu, Qiyang Zhou, Anning Hu, Li Shen, Ya Zhang, Dacheng Tao·
arXiv:2609.34426v2 Announce Type: replace Abstract: This paper addresses the problem of safe offline reinforcement learning, which involves training a policy to satisfy safety constraints using an offline dataset. This problem is inherently challenging as it requires balancing th…
arXiv:2609.36421v1 Announce Type: cross Abstract: Reinforcement learning (RL) is a powerful framework for robotic control, yet its practical application is often hindered by high sample complexity. This is particularly restrictive in physical domains where interaction data is cos…
arXiv cs.LG
TIER_1English(EN)·T. Konstantin Rusch, Tim Seyde, Jared Boyer, Zach J. Patterson, Daniela Rus·
arXiv:2609.37432v1 Announce Type: new Abstract: Looped reasoning models repeatedly apply a shared set of parameters, enabling more computation without increasing the model size. These models also support input-dependent computation by dynamically deciding when to stop looping. Mo…
arXiv:2609.36945v1 Announce Type: new Abstract: We study the learning dynamics of fine-tuning a policy model on self-generated and reward-weighted data, with particular focus on a generalized version of REINFORCE -- referred to as RE(S) -- that updates the rollout distribution on…
arXiv:2609.36864v1 Announce Type: new Abstract: Group-relative methods for reinforcement learning with verifiable rewards (RLVR) learn from differences in rollout outcomes. Independently sampling complete trajectories is costly and does not explicitly explore the decision space a…
arXiv:2609.36859v1 Announce Type: new Abstract: We identify and study a structural mechanism for Markovian nonconvex ADMM in reinforcement learning. Using finite discounted MDPs as a canonical proving ground, we show that the discounted Bellman resolvent $(I-\gamma P_\pi)^{-1}$ c…
arXiv:2609.36816v1 Announce Type: new Abstract: Reinforcement learning (RL) has become a cornerstone for improving the reasoning capabilities of large language models (LLMs), but the need for on-policy data substantially limits training efficiency. Reusing off-policy data through…
arXiv:2609.36486v1 Announce Type: new Abstract: We study an unknown-transition finite-horizon Markov decision process (MDP) with a finite collection of known reward functions $\{r^1, r^2, \ldots, r^M\}$. The goal is to output an $\epsilon$-optimal policy for every reward using on…
arXiv cs.LG
TIER_1English(EN)·Hsiao-Ru Pan, Florent Draye, Bernhard Sch\"olkopf·
arXiv:2609.36058v1 Announce Type: new Abstract: Recent progress in reinforcement learning with verifiable rewards (RLVR) has highlighted the effectiveness of simple critic-free policy-gradient methods such as Group Relative Policy Optimization (GRPO). In contrast, actor-critic me…
arXiv:2609.35880v1 Announce Type: new Abstract: Offline reinforcement learning (RL) has traditionally focused on learning policies for direct deployment under conservative objectives, where uncertainty outside the offline dataset is treated pessimistically to ensure robustness. W…
arXiv:2502.04492v3 Announce Type: replace Abstract: The advancement of LLMs and their accessibility have triggered renewed interest in multi-agent reinforcement learning as robust and adaptive frameworks for dynamically changing environments. This paper introduces \texttt{RL-Foca…
arXiv cs.AI
TIER_1English(EN)·Emiliano Penaloza, Dane Malenfant, Dheeraj Vattikonda, Roger Creus Castanyer, Siddarth Venkatraman, Abhay Puri, Jonathan Light, Matthew James Sargent, Augustine N. Mavor-Parker, Massimo Caccia, Lucas Caccia, Glen Berseth, Esmeralda S. Whitammer, Alessand…·
arXiv:2609.35750v2 Announce Type: replace-cross Abstract: Scaling the horizon of agentic LLMs is bottlenecked by the need to fit ever longer context traces in GPU memory. Context compaction has been the most popular mechanism to alleviate this issue, keeping GPU memory constant f…
arXiv cs.AI
TIER_1English(EN)·Minung Kim, Jeongmo Kim, Gwanwoo Choi, Seungyul Han·
arXiv:2605.13054v2 Announce Type: replace-cross Abstract: Cross-domain offline reinforcement learning learns a target policy from pre-collected source and target datasets with different dynamics. When target data are scarce, effectively compensating for their limited coverage usi…
arXiv cs.AI
TIER_1English(EN)·Xian Yu, Minheng Xiao, Lei Ying·
arXiv:2405.14749v3 Announce Type: replace-cross Abstract: Risk-sensitive reinforcement learning (RL) is crucial for maintaining reliable performance in high-stakes applications. While traditional RL methods aim to learn a point estimate of the random cumulative cost, distribution…
arXiv:2605.14483v2 Announce Type: replace Abstract: Large language models (LLMs) provide a flexible foundation for multi-agent systems, but their effectiveness and computational cost depend critically on orchestration design. Across different tasks, role design, capacity assignme…
arXiv cs.AI
TIER_1English(EN)·Hongyang Li, Xiao Li, Caesar Wu, Said Mammar, Gr\'egoire Danoy, Pascal Bouvry·
arXiv:2609.37119v1 Announce Type: cross Abstract: Recent approaches to reinforcement learning (RL) post-training for large language models increasingly remove the critic to reduce training instability and memory overhead. Even where a critic is trained, it is discarded once train…
arXiv:2609.36552v1 Announce Type: cross Abstract: Maximum Likelihood Reinforcement Learning (MaxRL) targets prompt-wise log-success and has shown strong performance on reasoning tasks. Under finite rollout budgets, however, the estimator used by MaxRL attenuates each prompt's lik…
arXiv:2609.36393v1 Announce Type: cross Abstract: Traditional reinforcement learning (RL) techniques focus on maximizing expected cumulative reward, where each action assumes to take a constant unit of time. However, this assumption does not hold for agentic RL tasks such as mach…
arXiv cs.AI
TIER_1English(EN)·Muhammad U. Nasir, Alex Vogt, Steven D. James, Julian Togelius·
arXiv:2609.36867v1 Announce Type: new Abstract: We propose STRAT, an auxiliary task that trains deep reinforcement learning (RL) agents to predict a short textual trace of their own state. Inspired by human spatial navigation, the description combines landmark, route, and survey …
arXiv cs.AI
TIER_1English(EN)·Zhenrui Yue, Huimin Zeng, Yueqi Wang, Yaokun Liu, Fengran Mo, Jinghan Zhang, Mung Yao Jia, Gyuseok Lee, Yang Zhang, Na Wei, Dong Wang·
arXiv:2609.36742v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for improving large language models (LLMs) on various tasks, yet its sparse outcome rewards lack token-level credit assignment for intermediate ste…
Most successes of superhuman game-playing algorithms are in games with discrete actions, yet in auctions, robotics, sports, or trading, actions are nearly continuous. Prior techniques either rely on expert-designed discretizations or are sample inefficient. We present a scalable …
arXiv cs.LG
TIER_1English(EN)·Yichen Lin, Xuyuan Xiong, Xue Wang, Xiangfu Meng, Mike Mingcheng Wei, Tao Yao·
arXiv:2609.30391v1 Announce Type: new Abstract: Existing in-context reinforcement learning methods mainly pretrain Transformers with supervised behavior-prediction objectives. This enables task inference from context, but makes the learned policy strongly depend on the quality of…
arXiv cs.AI
TIER_1English(EN)·Mattie Terzolo, Mikolaj Sacha, Ayan Sinha, Andrew Rabinovich·
arXiv:2609.27035v2 Announce Type: replace Abstract: Group Relative Policy Optimization (GRPO) and related policy-gradient methods for training language model agents collapse an entire multi-turn rollout into a single scalar trajectory reward before it enters the policy update. Wh…
Group-relative methods for reinforcement learning with verifiable rewards (RLVR) learn from differences in rollout outcomes. Independently sampling complete trajectories is costly and does not explicitly explore the decision space at critical positions. Feedback on a completed tr…
This paper addresses the problem of safe offline reinforcement learning, which involves training a policy to satisfy safety constraints using an offline dataset. This problem is inherently challenging as it requires balancing three highly interconnected and competing objectives: …
Robots that learn from a few demonstrations often require two forms of generalization. Compositional generalization recombines skills to solve new tasks, and skill generalization lets the learned policy behind each skill work in new situations. The two depend on each other, yet i…
Reinforcement learning (RL) is increasingly used to post-train vision-language-action (VLA) models, yet how RL reshapes these policies remains poorly understood. We find that RL across widely used flow-based VLA models, including π_{0.5} and GR00T~N1.5/N1.6, on LIBERO, ManiSkill,…
Test-time reinforcement learning (TTRL) enables models to improve their reasoning without relying on labeled training data, but existing approaches typically optimize a large fraction of the model parameters. This raises a natural question: can effective test-time adaptation emer…
Unified multimodal models can both look at and render images, so in principle they can repair their own generations: diagnose what an image gets wrong, revise it, observe the result, and diagnose again. Whether a revision helps is known only after it is rendered, so the reflectio…
Large language model (LLM) agents increasingly undertake extreme-long (xlong) horizon tasks, where a single execution can span hours, hundreds of model--environment interactions, and nearly 1M tokens per rollout. Applying online reinforcement learning (RL) to such executions pose…
A quadrilateral block decomposition of a planar domain is judged by whether it is complete, whether its elements are well shaped, and how many of its vertices are irregular. The last has a provable floor: the discrete Gauss-Bonnet identity enforces a lower bound on the total vert…
arXiv cs.AI
TIER_1English(EN)·James Wu, Chris R. Sims·
arXiv:2609.28737v1 Announce Type: cross Abstract: Biological agents do not learn under conditions of unlimited computation. For humans, learning and choice are shaped by constraints on perception, attention, and working memory, which limit how much state information guides behavi…
arXiv:2610.09890v1 Announce Type: cross Abstract: Inverse optimization estimates the weights of an objective function that explain observed decisions as optimal solutions, and is used in a variety of fields. For mixed-integer linear programs (MILPs), existing methods aim to repro…
arXiv:2610.03423v1 Announce Type: new Abstract: Reinforcement learning (RL) for Text-to-3D (T23D) generation requires optimization across multiple quality dimensions such as semantic alignment and texture clarity. Existing methods typically optimize these dimensions simultaneousl…
arXiv:2610.00758v1 Announce Type: cross Abstract: By learning transferable rewards, inverse reinforcement learning (IRL) enables counterfactual evaluation of agents under modified environments. Such transfer places strict requirements on coverage since target environments affect …
arXiv:2610.01973v1 Announce Type: new Abstract: Reinforcement learning (RL) for video generation usually assigns one scalar reward to an entire sampled video. Yet a video is not uniformly flawed: some visual tokens may already satisfy the prompt, whereas others require correction…
arXiv:2505.16741v5 Announce Type: replace-cross Abstract: Minimum attention applies the least action principle in changes of control concerning state and time, first proposed by Brockett. The involved regularization is highly relevant in emulating biological control, such as moto…
arXiv:2106.12886v3 Announce Type: replace-cross Abstract: Modern machine learning approaches to classification, including AdaBoost, support vector machines, and deep neural networks, utilize surrogate loss techniques to circumvent the computational complexity of minimizing empiri…
arXiv:2609.38383v1 Announce Type: cross Abstract: Random exploration reveals how an environment can be traversed before a goal is specified. Can this experience support long-range planning without policy-improvement training? Our random-walk analysis explains what temporal relati…
arXiv stat.ML
TIER_1English(EN)·Dylan J. Foster, Alexander Rakhlin·
arXiv:2312.16730v2 Announce Type: replace-cross Abstract: Interactive decision making is the problem of learning to act well in an unknown environment, using the data that one's own actions generate to continuously improve, and arises in situations ranging from online platforms a…
<figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*bzw_rrScey_iRtRBlLodOg.png" /></figure><p>This is the Part-1 of Series <strong>Reinforcement Learning from Scratch. </strong>In this series we plan to cover the core concepts of modern RL from Q-table, Deep-Q Lea…
Medium — Claude tag
TIER_1English(EN)·Arjun Keerthi·
<h1> LoGRA: How Low-Rank Gradient Sketches Make LLM Reinforcement Learning Fit on Real Hardware </h1> <p>Reinforcement learning post-training has become one of the most effective ways to improve large language model reasoning. Models like DeepSeek-R1 demonstrated that RL-based fi…
<p>Somebody handed me a list of reinforcement learning algorithms and asked whether any of them could help my tools. It is a ChatGPT session printed to PDF: roughly ninety algorithm names across fifteen sections, with no descriptions, no citations and no results. I want to be pre…