PulseAugur
中
实时 07:16:49
English(EN) KV-streams for Efficient Compaction in Agentic Reinforcement Learning

AI 研究在数学、适应性和公平性方面推进强化学习 · 跟踪 10 个来源

研究人员正在探索先进的强化学习技术,以提高 AI 的数学推理和适应能力。一篇论文介绍了函数结构图强化学习 (FSG-RL),通过从函数图生成代码并使用多验证器反馈来增强数学问题解决能力,显著提高了准确性。其他研究侧重于通过样本重用来优化近端策略优化 (PPO),并探索用于测试时适应的无数据方法。此外,还在开发无监督强化学习、具有神经进化的持续学习以及将公平性与 AI 系统的可解释性相结合的新方法。 AI

影响 强化学习技术的进步可能带来更强大的 AI,以应对复杂的推理、适应性和伦理考量。

排序理由 多篇 arXiv 论文发表,详细介绍了强化学习中的新颖方法和分析。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 149 个来源。 我们如何撰写摘要 →

AI 研究在数学、适应性和公平性方面推进强化学习 · 跟踪 10 个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
多篇 arXiv 论文发表,详细介绍了强化学习中的新颖方法和分析。
Source corroboration
149 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
13 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+65 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [149]

  1. arXiv cs.LG TIER_1 English(EN) · Ocan Sankur (DEVINE), Thierry J\'eron (DEVINE), Nicolas Markey (DEVINE), David Mentr\'e (MERCE-France), Reiya Noguchi ·

    基于需求的测试:用博弈论增强强化学习

    arXiv:2407.18994v2 Announce Type: replace-cross Abstract: We consider the automatic online synthesis of black-box test cases from functional requirements specified as automata for reactive implementations. The goal of the tester is to reach some given state, so as to satisfy a co…

  2. arXiv cs.CL TIER_1 English(EN) · Zirui Li, Rech Silas, Lauri Juvela, Tom Backstrom, Mikko Kurimo ·

    EmphTTS:一种具有强化学习的强调控制TTS

    arXiv:2609.27599v2 Announce Type: cross Abstract: Generating controllable and human-like emphasis remains an open challenge in text-to-speech, even when explicit emphasis control signals are provided in the text input, limiting the communicative accuracy of synthetic speech in re…

  3. arXiv cs.CL TIER_1 English(EN) · Yifan Zhang ·

    关于 KL 正则化策略优化

    arXiv:2610.08963v1 Announce Type: cross Abstract: Asynchronous reinforcement learning (RL) for large language model (LLM) agents trains one policy on trajectories generated by another: rollouts come from stale checkpoints, and the inference engine's probabilities differ from the …

  4. arXiv cs.CL TIER_1 English(EN) · Shaokun Zhang, Yifan Zhang, Jian Hu, Yueying Li, Hao Zhang, Binfeng Xu, Jan Kautz, Yi Dong ·

    LoGRA: 使用低秩梯度草图扩展 LLM 强化学习

    arXiv:2610.06647v2 Announce Type: replace Abstract: Reinforcement learning has greatly advanced the capabilities of large language models, but its memory demands remain a barrier to broader adoption. We introduce LoGRA, an approach to RL post-training that reduces memory by retai…

  5. arXiv cs.LG TIER_1 English(EN) · Shripad V. Deshmukh, Yaswanth Chittepu, Dhawal Gupta, Philip Thomas, Scott Niekum ·

    凸凹强化学习

    arXiv:2610.09108v1 Announce Type: new Abstract: Policy learning drives many of the most consequential and heavily-invested applications of reinforcement learning today. Yet the core optimization problem it rests on (maximizing expected return) is notoriously non-convex, even unde…

  6. arXiv cs.LG TIER_1 English(EN) · Michael Tang, Mahmoud Abdelgalil, Jorge I. Poveda ·

    欺骗性土匪问题:探索性耦合与多智能体学习的脆弱性

    arXiv:2610.09120v1 Announce Type: new Abstract: Randomized exploration is central to bandit learning, multi-agent reinforcement learning, and zeroth-order policy search, yet its independence and privacy are usually only treated as technical assumptions. We show that these propert…

  7. arXiv cs.LG TIER_1 English(EN) · John L. Zhou, Yuxuan Dong, Jonathan C. Kao ·

    目标条件策略学习中的地平线信息诅咒

    arXiv:2610.09247v1 Announce Type: new Abstract: The difficulty of learning goal-reaching policies is often attributed to a "curse of horizon" that manifests as bias accumulation in temporal-difference backups and noisy advantage estimates. In this work, we identify an additional …

  8. arXiv cs.LG TIER_1 English(EN) · Zicheng Hu, Zhijian Zhou, Xuan Zhang, Yuchen Liu, Cheng Chen, Yuan Li, Qi Gu, Yan Feng, Hongyan Hao, Chao Qu ·

    COPC:异步大语言模型强化学习的耦合离轨策略校正

    arXiv:2610.09597v1 Announce Type: new Abstract: Asynchronous RL accelerates large language model post-training by decoupling rollout generation from optimization, but trains on stale trajectories. Existing methods primarily correct token-level policy mismatch through importance-r…

  9. arXiv cs.LG TIER_1 English(EN) · Mahmoud Selim, Cristina Cipriani, Karl Henrik Johansson ·

    超越策略支持:交互受限的离线强化学习在自动驾驶中的应用

    arXiv:2610.09763v1 Announce Type: new Abstract: Offline reinforcement learning enables reward-driven policy improvement from fixed datasets without requiring online exploration, making it particularly attractive in safety-critical domains. A central challenge, however, is distrib…

  10. arXiv cs.LG TIER_1 English(EN) · Yongqiang Yao, Jinru Tan, Kaihuan Liang, Zixin Yin, Yazhe Niu, Ruihao Gong, Dahua Lin, Ningyi Xu ·

    RollVerify:在长尾部署强化学习中实现效率与准确性的结合

    arXiv:2610.09914v1 Announce Type: new Abstract: Reinforcement learning is crucial for improving large language models' reasoning and generalization. It relies on massive rollouts whose lengths become increasingly long-tailed as context windows grow. In on-policy training, these l…

  11. arXiv cs.LG TIER_1 English(EN) · Tommaso Marzi, Ahmed Hendawy, Jan Peters, Carlo D'Eramo, Andrea Cini, Cesare Alippi ·

    持续图多智能体强化学习

    arXiv:2610.10302v1 Announce Type: new Abstract: In Continual Multi-Agent Reinforcement Learning (CMARL), agents learn cooperative policies across sequences of tasks, aiming to adapt effectively to new tasks while preserving the ability to solve previously encountered ones. In man…

  12. arXiv cs.LG TIER_1 English(EN) · Huizhen Yu, Isaiah Heidt ·

    多链MDP的平均奖励强化学习:一种分层分解方法

    arXiv:2610.10326v1 Announce Type: new Abstract: We study learning optimal policies in average-reward multichain Markov decision processes (MDPs), where the optimal gain may depend on the initial state and recurrence structures vary across policies, creating challenges for reinfor…

  13. arXiv cs.LG TIER_1 English(EN) · Fernando Martinez, Abhishek Satyam, Tao Li, Junaid Farooq, Ying Wang, Juntao Chen ·

    专家访谈:LLM 引导的强化学习在自主网络防御中的应用

    arXiv:2610.09337v1 Announce Type: cross Abstract: Policy-based reinforcement learning (RL) approaches have produced promising results for autonomous cyber defense; however, they are sample-inefficient in settings where defenders must respond under delayed, partial observations wi…

  14. arXiv cs.LG TIER_1 English(EN) · Lijun Bo, Yijie Huang, Chenhao Lu ·

    面向超立方体状态空间的边界感知强化学习与确定性策略梯度

    arXiv:2610.09712v1 Announce Type: cross Abstract: We develop a continuous-time deterministic policy gradient framework for reinforcement learning with reflected state dynamics, where the state process is governed by a controlled reflected stochastic differential equation on a hyp…

  15. arXiv cs.LG TIER_1 English(EN) · Benedikt Koch, Winston Chou, Aur\'elien Bibaut, Nathan Kallus ·

    Policy Learning with Weak Signals

    arXiv:2610.10167v1 Announce Type: cross Abstract: Policy learning in digital experimentation faces three challenges: weak signal-to-noise ratios, rich covariate spaces, and massive data volumes. We formalize this regime by modeling treatment-effect estimates from increasingly fin…

  16. arXiv cs.LG TIER_1 English(EN) · Peng Xia, Jianwen Chen, Hanyang Wang, Jiaqi Liu, Kaide Zeng, Yu Wang, Siwei Han, Yiyang Zhou, Xujiang Zhao, Haifeng Chen, Zeyu Zheng, Cihang Xie, Huaxiu Yao ·

    SkillRL:通过递归技能增强强化学习实现智能体进化

    arXiv:2602.08234v2 Announce Type: replace Abstract: Large Language Model (LLM) agents have shown stunning results in complex tasks, yet they often operate in isolation, failing to learn from past experiences. Existing memory-based methods primarily store raw trajectories, which a…

  17. arXiv cs.AI TIER_1 English(EN) · Ange Tong ·

    学会何时优化:面向预算神经网络算子偏微分方程求解器的长时域强化学习

    arXiv:2610.06883v1 Announce Type: cross Abstract: Neural operators provide fast surrogates for time-dependent PDEs, but autoregressive deployment creates a refinement-allocation problem: prediction errors vary over space and time, while only a finite number of local corrections c…

  18. arXiv cs.LG TIER_1 English(EN) · Silong Yong, Stephen Sheng, Carl Qi, Xiaojie Wang, Evan Sheehan, Anurag Shivaprasad, Yaqi Xie, Katia Sycara, Yesh Dattatreya ·

    面向长时域机器人任务的可泛化稠密奖励

    arXiv:2604.00055v2 Announce Type: replace-cross Abstract: Existing robotic foundation policies are trained primarily via large-scale imitation learning. While such models demonstrate strong capabilities, they often struggle with long-horizon tasks due to distribution shift and er…

  19. arXiv cs.LG TIER_1 English(EN) · Valliappan Chidambaram Adaikkappan, Sai Rajeswar, Pietro Mazzaglia, David Meger ·

    MSPR:面向目标条件强化学习的多尺度预测表示

    arXiv:2605.09364v2 Announce Type: replace Abstract: This paper investigates robust representation learning in offline goal-conditioned reinforcement learning (GCRL). Particularly in sparse reward scenarios, learning representations that align state and goal latents is a challenge…

  20. arXiv cs.LG TIER_1 English(EN) · Junsoo Ha ·

    随机梯度下降上升法对于非凸PL极小极大博弈而言并非最优

    arXiv:2610.07814v1 Announce Type: cross Abstract: How far can stochastic gradient descent ascent (SGDA) go by tuning its timescale ratio and step sizes in nonconvex min-max games? We answer this question for nonconvex-PL (NC-PL) games by establishing the first tight complexity of…

  21. arXiv cs.LG TIER_1 English(EN) · Fernando Martinez, Tao Li, Yingdong Lu, Juntao Chen ·

    具有反事实语义-社会世界模型的独立多智能体强化学习

    arXiv:2610.07704v1 Announce Type: cross Abstract: Fully decentralized multi-agent reinforcement learning (MARL), also referred to as independent learning, requires each agent to learn and act using only its local information and experience, without a centralized critic or inter-a…

  22. arXiv cs.LG TIER_1 English(EN) · Naoki Nishikawa, Taiji Suzuki ·

    用于分层推理奖励的强化学习:Transformer 的 Minimax 最优速率

    arXiv:2610.08561v1 Announce Type: new Abstract: Reinforcement learning (RL) has become a standard tool for post-training language models on reasoning tasks, where the policy is updated by reward feedback while exploring the space of responses. Despite its empirical success, theor…

  23. arXiv cs.LG TIER_1 English(EN) · Fengxu Liu, Siwei Wang, Gal Dalal, Shie Mannor, Yihan Du ·

    线性函数逼近下的分段奖励反馈强化学习

    arXiv:2610.08271v1 Announce Type: new Abstract: Classical reinforcement learning (RL) assumes that a reward is observed for every visited state-action pair. However, in real-world applications such as autonomous driving, such fine-grained feedback can be costly or difficult to co…

  24. arXiv cs.LG TIER_1 English(EN) · SungJae Ahn, Jeong Woon Lee, Kyoleen Kwak, Hyoseok Hwang ·

    深度强化学习中平滑控制的时间正则化再探讨

    arXiv:2610.07910v1 Announce Type: new Abstract: Deep Reinforcement Learning policies can produce nonsmooth action oscillations that hinder deployment on physical robots. Existing architectural and penalty-based approaches seek spatial smoothness by directly reducing sensitivity t…

  25. arXiv cs.CL TIER_1 English(EN) · Bolian Li, Ting-Yao Hu, Cheng-Yu Hsieh, Sanjoy Chowdhury, Oncel Tuzel, Raviteja Vemulapalli ·

    为 Agentic 强化学习构建 MoE 专家选择结构

    arXiv:2610.07332v1 Announce Type: cross Abstract: Long-horizon LLM agents are frequently implemented using sparse mixture-of-experts (MoE) models, yet the co-design of agentic behavior and MoE structures remains underexplored. In this work, we comprehensively study the connection…

  26. arXiv cs.AI TIER_1 English(EN) · Chenliang Li, Adel Elmahdy, Alex Boyd, Zhongruo Wang, Siliang Zeng, Alfredo Garcia, Parminder Bhatia, Taha Kass-Hout, Cao Xiao, Mingyi Hong ·

    通过回合级重要性采样和剪辑触发归一化稳定长时域LLM智能体的离策略训练

    arXiv:2511.20718v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) algorithms such as PPO and GRPO are widely used to train large language models (LLMs) for multi-turn agentic tasks. However, in off-policy training pipelines, these methods can exhibit unstable …

  27. arXiv cs.AI TIER_1 English(EN) · Wenwen Si, Honghao Wei ·

    基于共形动作集的强化学习:在序列推荐中的应用

    arXiv:2610.08743v1 Announce Type: cross Abstract: Sequential recommenders typically use a fixed slate size even though the number of useful alternatives changes within a session. We propose Reinforcement Learning with Calibrated Pruning (RLCP), which adapts the retained action se…

  28. arXiv cs.AI TIER_1 English(EN) · Guhyeon Kang, Minhae Kwon ·

    方差规避型 $n$ 步离线强化学习用于稀疏长时序环境

    arXiv:2610.07899v1 Announce Type: cross Abstract: Generative actors are transforming offline reinforcement learning (RL) by enabling expressive policy classes that model complex action distributions. However, this expressiveness also exposes a key challenge in heterogeneous datas…

  29. arXiv cs.AI TIER_1 English(EN) · Qi Liu, Fengming Liang, Yiqun Chen, Erhan Zhang, Jiaxin Mao ·

    在嵌入空间中通过强化学习进行检索学习

    arXiv:2610.07731v1 Announce Type: cross Abstract: Dense retrieval models are typically trained with contrastive objectives that learn effective representations but do not directly optimize retrieval metrics or downstream task performance. To address this problem, we introduce REL…

  30. arXiv cs.AI TIER_1 English(EN) · Xiaoyang Cao, Jingqi Li, Zhe Fu, Alexandre M. Bayen ·

    谁来承担责任?多智能体强化学习中共享约束的责任学习

    arXiv:2610.07491v1 Announce Type: cross Abstract: When multiple agents share a cost budget, a common Lagrange multiplier can enforce the aggregate constraint but does not determine how its penalty should be allocated across agents. Uniform penalties ignore heterogeneity in the re…

  31. arXiv cs.AI TIER_1 English(EN) · T. Y. Tsui, Zihao Ye, Pengxiang Cai, Yanchao Li, Yuqiang Li, Zhehong Ai ·

    最小见证强化学习

    arXiv:2610.07226v1 Announce Type: cross Abstract: ``What are the irreducible conditions that are sufficient to produce an outcome?'' is one of the most common questions that recur across computation and science. Its answers, the minimal sufficient witnesses, are what we mean by e…

  32. arXiv cs.AI TIER_1 English(EN) · Zhen Li, Shuai Zhang, Yanggan Gu, Yiming Zhang, Yang Yu, Mingfa Feng, Congkai Xie, Shuang Yu, Junjie Lai, Hongxia Yang ·

    TRIAGE:原生 NVFP4 强化学习的方向感知不匹配稳定化

    arXiv:2610.07043v1 Announce Type: cross Abstract: Low-precision execution can substantially accelerate reinforcement learning (RL) for large language models, but discrepancies between learner and sampler execution can destabilize policy optimization. In this paper, we characteriz…

  33. Hugging Face Daily Papers TIER_1 English(EN) ·

    深度强化学习中平滑控制的时间正则化再探讨

    Deep Reinforcement Learning policies can produce nonsmooth action oscillations that hinder deployment on physical robots. Existing architectural and penalty-based approaches seek spatial smoothness by directly reducing sensitivity to changes in state inputs, but their broad const…

  34. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Jiaxin Mao ·

    在嵌入空间中通过强化学习进行检索学习

    Dense retrieval models are typically trained with contrastive objectives that learn effective representations but do not directly optimize retrieval metrics or downstream task performance. To address this problem, we introduce RELER (REinforcement LEarning for Retrieval), a reinf…

  35. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Juntao Chen ·

    具有反事实语义-社会世界模型的独立多智能体强化学习

    Fully decentralized multi-agent reinforcement learning (MARL), also referred to as independent learning, requires each agent to learn and act using only its local information and experience, without a centralized critic or inter-agent communication. Such a stringent information s…

  36. Hugging Face Daily Papers TIER_1 English(EN) ·

    关于 KL 正则化策略优化

    Asynchronous reinforcement learning (RL) for large language model (LLM) agents trains one policy on trajectories generated by another: rollouts come from stale checkpoints, and the inference engine's probabilities differ from the trainer's even at identical parameters. Standard r…

  37. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Alexandre M. Bayen ·

    谁来承担责任?多智能体强化学习中共享约束的责任学习

    When multiple agents share a cost budget, a common Lagrange multiplier can enforce the aggregate constraint but does not determine how its penalty should be allocated across agents. Uniform penalties ignore heterogeneity in the rewards agents sacrifice, while agent-specific multi…

  38. Hugging Face Daily Papers TIER_1 English(EN) ·

    ThunderSyncRL:无损加速智能体强化学习

    Language models are moving beyond generating answers to pursuing long-horizon goals in interactive environments. Post-training these agents requires long, heterogeneous trajectories, and synchronous systems leave learner engines idle until rollout and verification finish. To sque…

  39. arXiv cs.LG TIER_1 Deutsch(DE) · Jinwoo Kim, Shraddha Barke ·

    理解强化学习中的富集

    arXiv:2610.02846v1 Announce Type: new Abstract: When rewards are sparse, reinforcement learning with verifiable rewards (RLVR) often uses hints or intermediate guidance to generate more successful rollouts. This enrichment biases policy-gradient updates unless corrected via impor…

  40. arXiv cs.CL TIER_1 English(EN) · Xin Wang, Wenhao Wu, Menghao Zhang, Zhi Wang, Kun Shao, Jian Luan ·

    AdaStep:用于智能体强化学习的自适应步长信用加权

    arXiv:2610.03223v1 Announce Type: cross Abstract: Long-horizon LLM agents are typically trained with sparse outcome rewards, making trajectory-level objectives too coarse to distinguish the contribution of individual decisions. Step-level credit assignment provides finer-grained …

  41. arXiv cs.LG TIER_1 English(EN) · Ganghun Lee, Minji Kim, Minsu Lee, Byoung-Tak Zhang ·

    奖励膨胀:强化学习的健康刺激

    arXiv:2610.02545v1 Announce Type: new Abstract: Reward serves as the primary learning signal in reinforcement learning (RL). However, while reward magnitudes are typically held fixed throughout training, their temporal modulation remains underexplored. In this paper, we propose r…

  42. arXiv cs.AI TIER_1 English(EN) · Arip Asadulaev, Aladin Djuhera, Karim Salta, Holger Boche, Fakhri Karray, Martin Takac ·

    热带强化学习

    arXiv:2610.02478v1 Announce Type: new Abstract: Reinforcement learning for large language models typically maximizes expected return, adding up the probabilities of all successful trajectories. However, the classical sum formulation can only report how often the model policy succ…

  43. arXiv cs.AI TIER_1 English(EN) · Mingyu Park, Samyeul Noh, Hyun Myung, Donghwan Lee ·

    VIGOR:基于模型的强化学习中通过潜在空间一致性实现零样本视觉泛化

    arXiv:2610.02801v1 Announce Type: new Abstract: Model-based reinforcement learning (MBRL) achieves strong sample efficiency by planning within learned latent dynamics, yet its performance degrades substantially under unseen visual distractions such as background variations, light…

  44. arXiv cs.AI TIER_1 English(EN) · Yuxuan Fan, Jaehong Yoon ·

    MetaRubric:学习基于评分标准的强化学习奖励

    arXiv:2610.02824v1 Announce Type: new Abstract: Rubric-based reinforcement learning extends reward-driven optimization to open-ended tasks by assigning partial credit to individual response requirements. However, rubric judges can assign a high criterion score even when the infor…

  45. arXiv cs.AI TIER_1 English(EN) · Xinjie Liu, Ruihan Zhao, Anirban Chaudhuri, Cyrus Neary, Ufuk Topcu, David Fridovich-Keil ·

    多保真策略梯度稳定数据稀疏强化学习

    arXiv:2610.02505v1 Announce Type: cross Abstract: Policy gradient methods for on-policy reinforcement learning (RL) can become unstable when expensive, scarce target-domain data yield noisy gradient estimates. We address this challenge by complementing limited high-fidelity (HF) …

  46. arXiv cs.LG TIER_1 English(EN) · Juri Pfammatter, Kaixian Qu, Clemens Schwarke, Victor Klemm, Marco Hutter ·

    用于强化学习的双向Voronoi偏置探索课程

    arXiv:2610.03395v1 Announce Type: new Abstract: Long-horizon tasks with sparse rewards pose an exploration bottleneck for goal-conditioned reinforcement learning: a policy started from the initial state rarely reaches the goal and receives no learning signal. Reference motions, h…

  47. arXiv cs.LG TIER_1 English(EN) · Amit Thakur, Mukesh Singhal ·

    面向开放团队多智能体强化学习的周转正交信用分配

    arXiv:2610.02847v1 Announce Type: cross Abstract: Open-team multi-agent reinforcement learning studies cooperative systems in which agents may join, leave, or be replaced during an episode. In such settings, the team return changes both because agents choose useful actions and be…

  48. arXiv cs.LG TIER_1 English(EN) · Matthew Brun, Xu Andy Sun ·

    关于策略优化成功条件化的收敛性

    arXiv:2610.03642v1 Announce Type: new Abstract: Success conditioning is a strategy for improving decision-making policies in stochastic environments; it updates a policy by increasing the probability of taking actions that yield successful outcomes. Success conditioning is common…

  49. arXiv cs.LG TIER_1 English(EN) · Yudong Lin, Haoyuan Deng, Zhuoxuan Yuan, Zaijia Yang, Yuanjiang Xue, Ziwei Wang ·

    UniIntervene++:一种用于高效现实世界强化学习的自适应干预代理

    arXiv:2610.03620v1 Announce Type: new Abstract: Online reinforcement learning (RL) enables robot policies to improve through physical interaction, but the assistance they require changes as their competence evolves. Existing intervention strategies based on offline estimates or f…

  50. arXiv cs.AI TIER_1 English(EN) · Hongqiang Lin, Dongxu Zhang, Yiding Sun, Mingzhe Li, Ning Yang, Haijun Zhang ·

    Offline Policy Optimization with Posterior Sampling

    arXiv:2605.07393v2 Announce Type: replace Abstract: A fundamental challenge in model-based offline reinforcement learning (RL) lies in the trade-off between generalization and robustness against exploitation errors in out-of-distribution (OOD) regions. The key to resolving this t…

  51. arXiv cs.AI TIER_1 English(EN) · Jiaxin Zhang, Xiangyu Peng, Qinglin Chen, Yu Li, Hiroaki Hayashi, Chien-Sheng Wu ·

    前瞻性追溯:通过预测-现实差距进行自校准强化学习

    arXiv:2610.02740v1 Announce Type: cross Abstract: Reinforcement learning for long-horizon agents relies on purely retrospective training signals: credit is assigned only after observing environmental consequences, leaving the agent's belief at action time invisible to the gradien…

  52. arXiv cs.AI TIER_1 English(EN) · Saptarshi Nath, Inish M. D'Souza, Antonio Carta, Soheil Kolouri, Andrea Soltoggio ·

    如何在终身强化学习中查找和重用策略以实现持续适应

    arXiv:2610.03119v1 Announce Type: cross Abstract: In lifelong reinforcement learning, retaining previously learned policies is not sufficient for effective transfer to a new task. Useful knowledge may be distributed across several prior policies, and its relevance may change as t…

  53. arXiv cs.AI TIER_1 English(EN) · Guilhem Loussouarn, Nancy Nayak, Kin K. Leung ·

    分阶段强化学习的单一或多重策略?

    arXiv:2610.03475v1 Announce Type: cross Abstract: Many reinforcement-learning (RL) problems are non-stationary yet structured and can be decomposed into phases, each with its own transition probabilities and reward functions. When the phase sequence is known, the common solution …

  54. arXiv cs.AI TIER_1 English(EN) · Jiaxuan Luo, Xingguo Xu, Shanshan Wang, Yuhan Zhou, Zhen Zhang ·

    CriticHack:在机器人策略优化下评估视觉奖励

    arXiv:2610.02527v1 Announce Type: cross Abstract: Learned visual reward models are increasingly used to optimize robot policies, yet a reward model can score an execution that acts on the wrong object as highly as one that completes the task. We show that optimizing such a reward…

  55. Hugging Face Daily Papers TIER_1 English(EN) ·

    为 Agentic 强化学习构建 MoE 专家选择结构

    Long-horizon LLM agents are frequently implemented using sparse mixture-of-experts (MoE) models, yet the co-design of agentic behavior and MoE structures remains underexplored. In this work, we comprehensively study the connections between agentic post-training and MoE expert sel…

  56. Hugging Face Daily Papers TIER_1 English(EN) ·

    TRIAGE:原生 NVFP4 强化学习的方向感知不匹配稳定性

    Low-precision execution can substantially accelerate reinforcement learning (RL) for large language models, but discrepancies between learner and sampler execution can destabilize policy optimization. In this paper, we characterize the interaction between mismatch and the policy-…

  57. Hugging Face Daily Papers TIER_1 English(EN) ·

    LoGRA:通过低秩梯度草图扩展大型语言模型强化学习

    Reinforcement learning (RL) has greatly advanced the capabilities of large language models (LLMs), but its memory demands remain a barrier to broader adoption. We introduce LoGRA, an approach to RL post-training that reduces memory by retaining useful learning signals in low-rank…

  58. Hugging Face Daily Papers TIER_1 English(EN) ·

    最小见证强化学习

    ``What are the irreducible conditions that are sufficient to produce an outcome?'' is one of the most common questions that recur across computation and science. Its answers, the minimal sufficient witnesses, are what we mean by explanations, mechanisms and reasons. These problem…

  59. arXiv cs.LG TIER_1 English(EN) · Morgan Byrd, Maks Sorokin, Robert Wright, Sehoon Ha ·

    Reward as Observation: Learning Reward-Based Policies for Rapid Adaptation

    arXiv:2610.00729v1 Announce Type: new Abstract: This paper explores a reward-based policy to achieve zero-shot transfer between source and target environments with completely different observation spaces. While humans can demonstrate impressive adaptation capabilities, deep neura…

  60. arXiv cs.LG TIER_1 English(EN) · Eleni Nisioti, Andrea Cossu, Kathrin Korte, Sebastian Risi ·

    基于神经进化的持续强化学习

    arXiv:2610.01583v1 Announce Type: cross Abstract: Despite many studies about causes and remedies of plasticity loss in Reinforcement Learning (RL) under continual task changes, no RL method has yet consistently achieved a good balance between adaptation and forgetting. Here we tu…

  61. arXiv cs.LG TIER_1 English(EN) · Tristan Shah, Wooyoung Chung, Volodomyr Makarenko, Juan Wachs, Stas Tiomkin ·

    Bellman 遇上 Lyapunov:通过掌握混沌实现无监督强化学习

    arXiv:2610.02012v1 Announce Type: new Abstract: Reinforcement learning (RL) is a powerful paradigm for training agents, yet its success rests on domain expertise of human engineers who design informative reward signals for every new task. Unsupervised RL aims to reduce this engin…

  62. arXiv cs.LG TIER_1 English(EN) · Zihan Liu, Xurong Xie ·

    具有可执行验证器的函数结构强化学习用于数学推理

    arXiv:2610.01729v1 Announce Type: new Abstract: Algorithmic mathematical reasoning requires reliable decomposition, computation, and aggregation. Final-answer rewards provide limited guidance on intermediate errors, while successful execution does not guarantee mathematical corre…

  63. arXiv cs.LG TIER_1 English(EN) · Ryunyi Lee, Kangjun Noh, Somin Kim, Heedong Kim, Kyungwoo Song ·

    Range-GRPO:通过奖励区间间的成对关系进行策略优化

    arXiv:2610.01548v1 Announce Type: new Abstract: As the use of large language models (LLMs) expands, post-training has become increasingly important for adapting them to downstream tasks. However, obtaining reliable supervision remains costly, especially in domains without referen…

  64. arXiv cs.LG TIER_1 English(EN) · Alessandro Montenegro, Riccardo Venturelli, Marco Mussi, Matteo Papini, Alberto Maria Metelli ·

    在近端策略优化中重用过去样本:何时以及如何提供帮助?

    arXiv:2610.01399v1 Announce Type: new Abstract: Among on-policy deep reinforcement learning methods, Proximal Policy Optimization (PPO) has become the de facto standard, due to its consistently strong empirical performance across diverse application domains. However, on-policy me…

  65. arXiv cs.LG TIER_1 English(EN) · Zhanming Zhang, Vinoth Selvendran ·

    锐化后适应:无数据测试时强化学习的入口状态锐化

    arXiv:2610.00903v1 Announce Type: new Abstract: Test-time reinforcement learning (TTRL) adapts language models on unlabeled test problems using supervision derived from their own samples. This makes the checkpoint's \emph{entry state} consequential: a diffuse policy provides nois…

  66. arXiv cs.LG TIER_1 English(EN) · Bente Hinkenhuis, Seyed Sahand Mohammadi Ziabari, Ali Mohammed Mansoor Alsahag ·

    在多实例强化学习系统中整合公平性和可解释性

    arXiv:2610.00035v1 Announce Type: new Abstract: Predicting student performance from educational interaction data requires models that are both accurate and sufficiently transparent to support meaningful intervention, while demographic information introduces an additional risk of …

  67. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Mukesh Singhal ·

    面向开放团队多智能体强化学习的周转正交信用分配

    Open-team multi-agent reinforcement learning studies cooperative systems in which agents may join, leave, or be replaced during an episode. In such settings, the team return changes both because agents choose useful actions and because the active population itself changes. Standa…

  68. arXiv cs.AI TIER_1 English(EN) · Jinhwan Sul, Alex Oshin, Evangelos A. Theodorou ·

    强化学习加速线性规划的对偶混合梯度

    arXiv:2610.01546v1 Announce Type: cross Abstract: Primal-dual hybrid gradient (PDHG) methods solve large-scale linear programs (LPs) using GPU-friendly matrix-vector products and projections, but their practical performance depends on coordinating algorithm parameters, accelerati…

  69. arXiv cs.AI TIER_1 English(EN) · Pedro Robles Dutenhefner, Dikshant Shehmar, Wagner Meira Jr., Marlos C. Machado ·

    为目标条件强化学习学习多个时间尺度

    arXiv:2610.00849v1 Announce Type: new Abstract: Existing approaches to offline goal-conditioned reinforcement learning (GCRL) struggle with long-horizon tasks. Discounting shrinks value differences between distant states until they fall below the function approximation error, lea…

  70. arXiv cs.AI TIER_1 English(EN) · Ziyi Chen, Yan Zhang, Jianhui Wei, Daoan Zhang, Zuozhu Liu ·

    面向Agentic强化学习的依赖感知奖励塑造

    arXiv:2610.01207v1 Announce Type: new Abstract: When training large language models with reinforcement learning, terminal rewards provide little guidance about which steps matter. Common methods for assigning step credit overlook that work built on uncorrected mistakes is wasted …

  71. arXiv cs.AI TIER_1 English(EN) · Shiu-Hong Kao, Yubo Zhao, Zhenyu Tian, Pengzhan Sun, Yicong Li, Angela Yao ·

    重新思考基于后验集中度的概率强化学习

    arXiv:2610.01458v1 Announce Type: new Abstract: Verifier-free reinforcement learning with probability-based rewards offers a promising way to train LLMs on general reasoning tasks where external verifiers are unavailable. Yet the reliability of these rewards, especially in long-h…

  72. arXiv cs.AI TIER_1 English(EN) · Jude Waide, Robert Lieck ·

    Decision Titan:离线强化学习中用于长期记忆的测试时训练

    arXiv:2610.01513v1 Announce Type: new Abstract: Long-term dependencies remain a major challenge for sequential decision-making in the field of AI: RNNs suffer from vanishing gradients and the limited expressivity of vector-based hidden states, whilst Transformer-based models are …

  73. arXiv cs.AI TIER_1 English(EN) · Abid Mohamed Nadhir, Ahmad Al Hanbali, Beggas Mounir ·

    同态优势算子:在全同态加密约束下稳定强化学习

    arXiv:2610.02074v1 Announce Type: new Abstract: Privacy-preserving machine learning presents significant deployment challenges on the cloud for intelligent systems with confidential data. Fully Homomorphic Encryption (FHE) offers a compelling solution for secure computation, pres…

  74. arXiv cs.AI TIER_1 English(EN) · Bo-Wen Zhang, Junwei He, Maoqi Liu, Feiran Li, Song-Lin Lv, Wentao Ma, Rongyi Lin, Shuhan Zhong, Lan-Zhe Guo ·

    T2SPO:用于智能体强化学习的轨迹到步策略优化

    arXiv:2610.00388v1 Announce Type: cross Abstract: Reinforcement learning enables large language model (LLM) agents to learn multi-step behaviors through interaction with their environments. However, rewards in many interactive tasks reflect only the final outcome, providing limit…

  75. arXiv cs.AI TIER_1 English(EN) · Oleg Shchendrigin, Egor Cherepanov, Aleksandr I. Panov, Alexey K. Kovalev ·

    ALER:强化学习的自适应可学经验重写

    arXiv:2610.00592v1 Announce Type: cross Abstract: In partially observable reinforcement learning (RL), a later observation can make stored information obsolete or change what it implies for the next decision. Memory architectures and benchmarks for RL mostly test retention, the a…

  76. arXiv cs.AI TIER_1 English(EN) · Christos Spyridon Koulouris, Carlo Campajola ·

    超越超竞争性结果:深度强化学习在最优执行博弈中的合谋行为

    arXiv:2610.00619v1 Announce Type: cross Abstract: In this paper, we extend earlier findings of supra-competitive outcomes in optimal-execution games by identifying a learned punitive mechanism that deters deviations and provides behavioural evidence of collusion. We investigate t…

  77. arXiv cs.AI TIER_1 English(EN) · Yujie Zhu, Charles A. Hepburn, Matthew Thorpe, Giovanni Montana ·

    最优传输与强化学习的结合:一篇综述

    arXiv:2610.01413v1 Announce Type: cross Abstract: Reinforcement learning (RL) algorithms frequently compare probability distributions, such as state visitation distributions induced by policies and experts, action distributions from learned policies and offline datasets, or trans…

  78. arXiv cs.AI TIER_1 English(EN) · Yafei Zhang, Songshuo Lu, Sicong Liao, Zhi Chen, Yaohua Tang ·

    CARM:面向LLM强化学习的感知取消响应掩码

    arXiv:2610.02039v1 Announce Type: cross Abstract: Recent years have witnessed the rapid adoption of reinforcement learning (RL) in large language model (LLM) post-training, with substantial gains in mathematical reasoning and code generation. In practical systems, however, policy…

  79. arXiv cs.AI TIER_1 English(EN) · Sebastian Sanokowski, Alireza Sarmadi, Majid Khadiv ·

    FERPO:前向熵正则化策略优化

    arXiv:2610.02198v1 Announce Type: cross Abstract: Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic. However, critics are typically trained to predict returns, and accurate value pr…

  80. arXiv cs.AI TIER_1 English(EN) · Debosmita Bhaumik, Julian Togelius, Georgios N. Yannakakis, Ahmed Khalifa ·

    为强化学习内容生成器学习局部约束

    arXiv:2605.13570v2 Announce Type: replace Abstract: Constraint-based game content generators that learn local constraints from existing content, such as Wave Function Collapse (WFC), can generate visually satisfying game levels but face challenges in optimizing global properties,…

  81. arXiv cs.AI TIER_1 English(EN) · Yunho Choi, Jongwon Lim, Woojin Ahn, Minjae Oh, Jeonghoon Shim, Yohan Jo ·

    你的语言模型是它自己的批评者:基于演员内部状态的价值估计强化学习

    arXiv:2605.07579v3 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) for Large Reasoning Models rests on variance reduction, which requires both a reliable baseline and high prompt diversity within each training batch. This is especially…

  82. arXiv cs.LG TIER_1 English(EN) · Sathya Kamesh Bhethanabhotla, Efstratios Gavves, Andr\'e Biedenkapp ·

    具有复杂(值)记忆的强化学习

    arXiv:2609.38598v1 Announce Type: new Abstract: Partially observable environments pose a fundamental challenge in deep reinforcement learning, requiring agents to compress temporal information from observations and maintain a memory to make effective decisions. While there exist …

  83. arXiv cs.LG TIER_1 English(EN) · Mintae Kim, Koushil Sreenath ·

    模型驱动的离线强化学习中的分布内想象力

    arXiv:2609.38673v1 Announce Type: new Abstract: Model-based offline reinforcement learning (MBORL) improves sample efficiency through model-generated trajectories. However, accumulative model error can drive imagined trajectories outside the offline data distribution, leading to …

  84. arXiv cs.LG TIER_1 English(EN) · Xinyi Ni, Lifeng Lai ·

    鲁棒的风险敏感型从损坏的人类反馈中进行强化学习

    arXiv:2609.38938v1 Announce Type: new Abstract: Reinforcement learning with human feedback (RLHF) learns from human comparisons, which can be corrupted or deliberately manipulated. This paper studies online risk-sensitive RLHF with static conditional value-at-risk (CVaR) under ad…

  85. arXiv cs.LG TIER_1 English(EN) · Kihyun Yu, Seoungbin Bae, Dabeen Lee ·

    通过状态增强学习无限视界平均奖励CMDPs

    arXiv:2609.39093v1 Announce Type: new Abstract: We study infinite-horizon average-reward constrained Markov decision processes (CMDPs) under the weakly communicating assumption. Existing high-probability guarantees for this setting either require computationally inefficient algor…

  86. arXiv cs.LG TIER_1 English(EN) · Shuxiao Xie, Shuyang Xie, Dezhi Ran, Wei Yang, Tao Xie ·

    向后状态策略是学习算法的一部分

    arXiv:2609.39813v1 Announce Type: new Abstract: Low-precision training rounds tensors that the backward pass reads again, often for several gradients; each use can read the forward's rounded value, the original, or a new random rounding. This backward-state policy looks like a me…

  87. arXiv cs.LG TIER_1 English(EN) · Seonvin Cho, Soohyun Choi, Songnam Hong ·

    面向离线强化学习的角色自适应策略优化

    arXiv:2609.40149v1 Announce Type: new Abstract: Policy regularization in offline reinforcement learning balances policy improvement against reliance on uncertain value estimates. This balance can differ between selecting actions for execution and supplying actions for critic boot…

  88. arXiv cs.LG TIER_1 English(EN) · Buse Y{\i}lmaz ·

    用于 SpTRSV 优化的强化学习引导图变换

    arXiv:2609.40159v1 Announce Type: cross Abstract: Sparse triangular solve (SpTRSV) is a fundamental kernel in numerous scientific and engineering applications. However, the data dependencies inherent in sparse triangular matrices significantly limit the available parallelism and …

  89. arXiv cs.LG TIER_1 English(EN) · Shuai Dong, Yongfu Zhu, Yuqi Xu, Weichu Xie, Liuwenpu, Ziyue Wang, Kaiwen Tuo, Congcong Wang, Siyuan Wang, Wenqi Shao, Shuai Yang, Ji Zhao, Caoyuan Ma, Wenzheng Chang, Taiqiang Wu, Xinlei Yu, Hongrui Wu, Xiaoxuan He, Fangke Chen, Dianyi Wang, Kanghui Ti… ·

    RL始于RL:关于策略蒸馏以改进强化学习

    arXiv:2609.28145v2 Announce Type: replace Abstract: Reinforcement learning (RL) improves reasoning, but its performance depends on the policy from which training begins. We study on-policy distillation (OPD) as a preparation stage for RL and ask whether its benefits extend beyond…

  90. Hugging Face Daily Papers TIER_1 English(EN) ·

    MetaRubric:学习基于评分标准的强化学习奖励

    Rubric-based reinforcement learning extends reward-driven optimization to open-ended tasks by assigning partial credit to individual response requirements. However, rubric judges can assign a high criterion score even when the information or action it requires is absent from the …

  91. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Qilei Li ·

    合作学习后:多智能体强化学习中的梯度路由与依赖优化器的维护

    Cooperative MARL is commonly evaluated through cooperation discovery from random initialization, leaving open whether continued optimization can destabilize learned cooperation. Actor-critic comparisons can also conflate critic presence with value gradients entering shared actor …

  92. arXiv cs.NE (Neural & Evolutionary) TIER_1 English(EN) · Sebastian Risi ·

    基于神经进化的持续强化学习

    Despite many studies about causes and remedies of plasticity loss in Reinforcement Learning (RL) under continual task changes, no RL method has yet consistently achieved a good balance between adaptation and forgetting. Here we turn to an alternative optimization paradigm, neuroe…

  93. arXiv cs.NE (Neural & Evolutionary) TIER_1 English(EN) · Sebastian Risi ·

    基于神经进化的持续强化学习

    Despite many studies about causes and remedies of plasticity loss in Reinforcement Learning (RL) under continual task changes, no RL method has yet consistently achieved a good balance between adaptation and forgetting. Here we turn to an alternative optimization paradigm, neuroe…

  94. arXiv cs.AI TIER_1 English(EN) · Yang chen, Yitan Zhang, Michael Witbrock, Shuyue Hu ·

    通过基于Q分数匹配的序列值恢复实现无环逆强化学习

    arXiv:2609.38955v1 Announce Type: cross Abstract: Inverse Reinforcement Learning (IRL) aims to recover a reward function that explains expert demonstrations. Existing IRL methods typically rely on a bi-level optimization procedure that alternates between reward learning and polic…

  95. arXiv cs.AI TIER_1 English(EN) · Yitong Qiao, Tiantian He, Lei Liu, Yue Shen, Jian Wang, Jinjie Gu, Zhixuan Chu ·

    从模仿到奖励发现:基于策略的代理强化学习预热

    arXiv:2609.39436v1 Announce Type: cross Abstract: Reinforcement learning with a verifiable reward (RLVR) offers a scalable approach to training language-model agents, yet sparse outcome rewards can leave early training with little signal for policy improvement. We identify an On-…

  96. arXiv cs.CL TIER_1 English(EN) · Junyu Lu, Shichao Weng, Zhiqiang Wang, Haojie Luo, Jingfan Zhang, Yuhua Zhou, Cheng Du, Yuzhuo Zhang, Xi Li, Jinwei Du, Tiancheng Feng, Chuan Xiao, Shuyuan Zheng ·

    OPTS-TTPO:通过树搜索增强有限样本策略梯度学习

    arXiv:2609.40035v1 Announce Type: new Abstract: The policy-gradient theorem gives the exact gradient under the current policy, but finite on-policy samples may miss rare high-return trajectories. We study whether tree search improves their coverage within a fixed budget while con…

  97. arXiv cs.AI TIER_1 English(EN) · Yifan Zhang, Qifan Zhang, Liang Zheng ·

    面向多线路公交车调度,完成感知的跨保真离线到在线强化学习

    arXiv:2609.39868v1 Announce Type: new Abstract: Exploratory reinforcement learning (RL) on an operating bus fleet is impractical,while policies trained only from historical data cannot acquire new experience. Hybrid Offline-and-Online (H2O) RL combines fixed target replay with si…

  98. arXiv cs.AI TIER_1 English(EN) · Ali Aouad, Aymane El Gadarri, Vivek F. Farias ·

    重新审视奖励优化的缩放定律

    arXiv:2609.38526v1 Announce Type: cross Abstract: Scaling laws for optimization against reward models in AI alignment have pinned down how performance depends on optimization effort---measured by a KL-divergence budget relative to a reference policy. Beyond a certain budget, over…

  99. arXiv cs.AI TIER_1 English(EN) · Seokmin Ko, Taewon Goo, Kihyuk Hong ·

    DAMPER:面向平滑策略的优先返回梯度控制

    arXiv:2609.38903v1 Announce Type: cross Abstract: Actor-critic methods achieve strong performance in continuous control, but their policies can produce highly oscillatory actions. A common remedy is to add auxiliary smoothness losses. However, their contribution can be negligible…

  100. arXiv cs.AI TIER_1 English(EN) · Yilie Huang, Wenpin Tang, Xunyu Zhou ·

    ART for Diffusion Sampling: 一种基于强化学习的时间步长调度方法

    arXiv:2601.18681v3 Announce Type: replace-cross Abstract: We consider time discretization for score-based diffusion models to generate samples from a learned reverse-time dynamic on a finite grid. Uniform and hand-crafted grids can be suboptimal given a budget on the number of ti…

  101. arXiv cs.AI TIER_1 English(EN) · Tianle Zhong, Neiwen Ling, Yifan Pi, Zijun Wei, Tianshu Yu, Geoffrey Fox, Peng Wu, Xiao Yu ·

    通过零失配参考诊断LLM强化学习中的训练推理失配

    arXiv:2605.14220v2 Announce Type: replace-cross Abstract: Modern LLM RL systems separate rollout generation from policy optimization. These two stages are expected to produce token probabilities that match exactly. However, implementation differences can make them assign differen…

  102. Hugging Face Daily Papers TIER_1 English(EN) ·

    用于 SpTRSV 优化的强化学习引导图变换

    Sparse triangular solve (SpTRSV) is a fundamental kernel in numerous scientific and engineering applications. However, the data dependencies inherent in sparse triangular matrices significantly limit the available parallelism and make efficient workload distribution challenging. …

  103. Hugging Face Daily Papers TIER_1 English(EN) ·

    通过状态增强学习无限视界平均奖励CMDPs

    We study infinite-horizon average-reward constrained Markov decision processes (CMDPs) under the weakly communicating assumption. Existing high-probability guarantees for this setting either require computationally inefficient algorithms or have suboptimal dependence on the numbe…

  104. arXiv cs.LG TIER_1 English(EN) · Sanghyun Hahn, Jonghyun Choi ·

    带反馈校正的动作分块近端策略优化

    arXiv:2609.36250v1 Announce Type: new Abstract: Action chunking provides temporal abstraction in reinforcement learning by selecting short action sequences instead of individual actions, but many existing approaches face two limitations in high-dimensional robotic control. First,…

  105. arXiv cs.LG TIER_1 English(EN) · Shengchao Hu, Peng Wang, Jifeng Hu, Qiyang Zhou, Anning Hu, Li Shen, Ya Zhang, Dacheng Tao ·

    用于安全离线强化学习的Q学习惩罚Transformer

    arXiv:2609.34426v2 Announce Type: replace Abstract: This paper addresses the problem of safe offline reinforcement learning, which involves training a policy to satisfy safety constraints using an offline dataset. This problem is inherently challenging as it requires balancing th…

  106. arXiv cs.LG TIER_1 English(EN) · Nida Zamir, I-Hong Hou ·

    具有个体惩罚约束的躁动土匪:近最优索引与深度强化学习

    arXiv:2604.04101v4 Announce Type: replace Abstract: This paper investigates the Restless Multi-Armed Bandit (RMAB) framework under individual penalty constraints to address resource allocation challenges in dynamic wireless networked environments. Unlike conventional RMAB models,…

  107. arXiv cs.LG TIER_1 English(EN) · Rayan Mazouz, Haibo Zhao, Chris Hillar, Christian Shewmake ·

    Equivariant 强化学习的样本复杂度

    arXiv:2609.36421v1 Announce Type: cross Abstract: Reinforcement learning (RL) is a powerful framework for robotic control, yet its practical application is often hindered by high sample complexity. This is particularly restrictive in physical domains where interaction data is cos…

  108. arXiv cs.LG TIER_1 English(EN) · T. Konstantin Rusch, Tim Seyde, Jared Boyer, Zach J. Patterson, Daniela Rus ·

    Looped Actor: 深度循环推理模型用于强化学习

    arXiv:2609.37432v1 Announce Type: new Abstract: Looped reasoning models repeatedly apply a shared set of parameters, enabling more computation without increasing the model size. These models also support input-dependent computation by dynamically deciding when to stop looping. Mo…

  109. arXiv cs.LG TIER_1 English(EN) · Zhiwei Wang, Yanxi Chen, Yaliang Li, Bolin Ding ·

    在自生成和奖励加权数据上进行微调:学习动态、收敛速度和离策略性的优势

    arXiv:2609.36945v1 Announce Type: new Abstract: We study the learning dynamics of fine-tuning a policy model on self-generated and reward-weighted data, with particular focus on a generalized version of REINFORCE -- referred to as RE(S) -- that updates the rollout distribution on…

  110. arXiv cs.LG TIER_1 English(EN) · Fanchao Chen, Hengyu Fu, Shivaram Venkataraman, Jiantao Jiao ·

    模型改变主意的地方:事后诸葛亮式分歧定位,用于可验证奖励的高效强化学习

    arXiv:2609.36864v1 Announce Type: new Abstract: Group-relative methods for reinforcement learning with verifiable rewards (RLVR) learn from differences in rollout outcomes. Independently sampling complete trajectories is costly and does not explicitly explore the decision space a…

  111. arXiv cs.LG TIER_1 English(EN) · Zhaojun Peng ·

    用于强化学习的马尔可夫非凸ADMM:超越光滑块的贝尔曼-解析器稳定性

    arXiv:2609.36859v1 Announce Type: new Abstract: We identify and study a structural mechanism for Markovian nonconvex ADMM in reinforcement learning. Using finite discounted MDPs as a canonical proving ground, we show that the discounted Bellman resolvent $(I-\gamma P_\pi)^{-1}$ c…

  112. arXiv cs.LG TIER_1 English(EN) · Ruichuan Huang, Jinghan Liu, Congliang Chen ·

    迈向更好的训练信号:优势裁剪策略优化

    arXiv:2609.36816v1 Announce Type: new Abstract: Reinforcement learning (RL) has become a cornerstone for improving the reasoning capabilities of large language models (LLMs), but the need for on-policy data substantially limits training efficiency. Reusing off-policy data through…

  113. arXiv cs.LG TIER_1 English(EN) · Zijun Chen, Zihan Zhang ·

    最优多奖励强化学习

    arXiv:2609.36486v1 Announce Type: new Abstract: We study an unknown-transition finite-horizon Markov decision process (MDP) with a finite collection of known reward functions $\{r^1, r^2, \ldots, r^M\}$. The goal is to output an $\epsilon$-optimal policy for every reward using on…

  114. arXiv cs.LG TIER_1 English(EN) · Hsiao-Ru Pan, Florent Draye, Bernhard Sch\"olkopf ·

    ABC:基于优势的强化学习可验证奖励控制变量

    arXiv:2609.36058v1 Announce Type: new Abstract: Recent progress in reinforcement learning with verifiable rewards (RLVR) has highlighted the effectiveness of simple critic-free policy-gradient methods such as Group Relative Policy Optimization (GRPO). In contrast, actor-critic me…

  115. arXiv cs.LG TIER_1 English(EN) · Tianwei Ni, Vineet Jain, Akash Karthikeyan, Pierre-Luc Bacon ·

    从静态策略到离线强化学习中的自适应先验

    arXiv:2609.35880v1 Announce Type: new Abstract: Offline reinforcement learning (RL) has traditionally focused on learning policies for direct deployment under conservative objectives, where uncertainty outside the offline dataset is treated pessimistically to ensure robustness. W…

  116. arXiv cs.CL TIER_1 English(EN) · Selim Furkan Tekin, Gaowen Liu, Ramana Rao Kompella, Ling Liu ·

    使用两阶段强化学习代理对大型语言模型集成进行动态优化

    arXiv:2502.04492v3 Announce Type: replace Abstract: The advancement of LLMs and their accessibility have triggered renewed interest in multi-agent reinforcement learning as robust and adaptive frameworks for dynamically changing environments. This paper introduces \texttt{RL-Foca…

  117. arXiv cs.AI TIER_1 English(EN) · Emiliano Penaloza, Dane Malenfant, Dheeraj Vattikonda, Roger Creus Castanyer, Siddarth Venkatraman, Abhay Puri, Jonathan Light, Matthew James Sargent, Augustine N. Mavor-Parker, Massimo Caccia, Lucas Caccia, Glen Berseth, Esmeralda S. Whitammer, Alessand… ·

    用于 Agentic 强化学习中高效压缩的 KV-streams

    arXiv:2609.35750v2 Announce Type: replace-cross Abstract: Scaling the horizon of agentic LLMs is bottlenecked by the need to fit ever longer context traces in GPU memory. Context compaction has been the most popular mechanism to alleviate this issue, keeping GPU memory constant f…

  118. arXiv cs.AI TIER_1 English(EN) · Minung Kim, Jeongmo Kim, Gwanwoo Choi, Seungyul Han ·

    生成式支持跨域离线强化学习的重新调整

    arXiv:2605.13054v2 Announce Type: replace-cross Abstract: Cross-domain offline reinforcement learning learns a target policy from pre-collected source and target datasets with different dynamics. When target data are scarce, effectively compensating for their limited coverage usi…

  119. arXiv cs.AI TIER_1 English(EN) · Xian Yu, Minheng Xiao, Lei Ying ·

    关于风险敏感强化学习的分布策略梯度算法的近似与收敛性

    arXiv:2405.14749v3 Announce Type: replace-cross Abstract: Risk-sensitive reinforcement learning (RL) is crucial for maintaining reliable performance in high-stakes applications. While traditional RL methods aim to learn a point estimate of the random cumulative cost, distribution…

  120. arXiv cs.AI TIER_1 English(EN) · Xudong Chen, Yixin Liu, Hua Wei, Kaize Ding ·

    COLLATOR:基于反事实强化学习的组合式多智能体编排

    arXiv:2605.14483v2 Announce Type: replace Abstract: Large language models (LLMs) provide a flexible foundation for multi-agent systems, but their effectiveness and computational cost depend critically on orchestration design. Across different tasks, role design, capacity assignme…

  121. arXiv cs.AI TIER_1 English(EN) · Hongyang Li, Xiao Li, Caesar Wu, Said Mammar, Gr\'egoire Danoy, Pascal Bouvry ·

    解锁批评者:LLM 训练后无奖励策略优化

    arXiv:2609.37119v1 Announce Type: cross Abstract: Recent approaches to reinforcement learning (RL) post-training for large language models increasingly remove the critic to reduce training instability and memory overhead. Even where a critic is trained, it is discarded once train…

  122. arXiv cs.AI TIER_1 English(EN) · Zihao Chen, Fanxiang Xiong, Hongran Ren, Xuefeng Bai, Zhongxiang Dai, Kehai Chen, Zhiguo Zhang, Zhiyong Wang, Yu Cheng ·

    SERA:用于最大似然强化学习的尺度均衡部署分配

    arXiv:2609.36552v1 Announce Type: cross Abstract: Maximum Likelihood Reinforcement Learning (MaxRL) targets prompt-wise log-success and has shown strong performance on reasoning tasks. Under finite rollout budgets, however, the estimator used by MaxRL attenuates each prompt's lik…

  123. arXiv cs.AI TIER_1 English(EN) · Muhang Tian, Sherry Yang ·

    Reward-rate Policy Gradient for Efficient Machine Learning Engineering Agents

    arXiv:2609.36393v1 Announce Type: cross Abstract: Traditional reinforcement learning (RL) techniques focus on maximizing expected cumulative reward, where each action assumes to take a constant unit of time. However, this assumption does not hold for agentic RL tasks such as mach…

  124. arXiv cs.AI TIER_1 English(EN) · Muhammad U. Nasir, Alex Vogt, Steven D. James, Julian Togelius ·

    强化学习中状态轨迹推理作为辅助任务

    arXiv:2609.36867v1 Announce Type: new Abstract: We propose STRAT, an auxiliary task that trains deep reinforcement learning (RL) agents to predict a short textual trace of their own state. Inspired by human spatial navigation, the description combines landmark, route, and survey …

  125. arXiv cs.AI TIER_1 English(EN) · Zhenrui Yue, Huimin Zeng, Yueqi Wang, Yaokun Liu, Fengran Mo, Jinghan Zhang, Mung Yao Jia, Gyuseok Lee, Yang Zhang, Na Wei, Dong Wang ·

    SIPO:将强化学习与同策略自蒸馏相结合

    arXiv:2609.36742v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for improving large language models (LLMs) on various tasks, yet its sparse outcome rewards lack token-level credit assignment for intermediate ste…

  126. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Tuomas Sandholm ·

    具有连续动作游戏的学习高斯混合的正则化策略梯度

    Most successes of superhuman game-playing algorithms are in games with discrete actions, yet in auctions, robotics, sports, or trading, actions are nearly continuous. Prior techniques either rely on expert-designed discretizations or are sample inefficient. We present a scalable …

  127. arXiv cs.LG TIER_1 English(EN) · Yichen Lin, Xuyuan Xiong, Xue Wang, Xiangfu Meng, Mike Mingcheng Wei, Tao Yao ·

    从弱数据到强策略:Q-Targets赋能可证明的上下文强化学习

    arXiv:2609.30391v1 Announce Type: new Abstract: Existing in-context reinforcement learning methods mainly pretrain Transformers with supervised behavior-prediction objectives. This enables task inference from context, but makes the learned policy strongly depend on the quality of…

  128. arXiv cs.AI TIER_1 English(EN) · Mattie Terzolo, Mikolaj Sacha, Ayan Sinha, Andrew Rabinovich ·

    基于分解子任务的强化学习

    arXiv:2609.27035v2 Announce Type: replace Abstract: Group Relative Policy Optimization (GRPO) and related policy-gradient methods for training language model agents collapse an entire multi-turn rollout into a single scalar trajectory reward before it enters the policy update. Wh…

  129. Hugging Face Daily Papers TIER_1 English(EN) ·

    模型改变主意的地方:用于可验证奖励的高效强化学习的滞后发散定位

    Group-relative methods for reinforcement learning with verifiable rewards (RLVR) learn from differences in rollout outcomes. Independently sampling complete trajectories is costly and does not explicitly explore the decision space at critical positions. Feedback on a completed tr…

  130. Hugging Face Daily Papers TIER_1 English(EN) ·

    用于安全离线强化学习的Q学习惩罚Transformer

    This paper addresses the problem of safe offline reinforcement learning, which involves training a policy to satisfy safety constraints using an offline dataset. This problem is inherently challenging as it requires balancing three highly interconnected and competing objectives: …

  131. Hugging Face Daily Papers TIER_1 English(EN) ·

    Agent Priors-guided Policy Learning

    Robots that learn from a few demonstrations often require two forms of generalization. Compositional generalization recombines skills to solve new tasks, and skill generalization lets the learned policy behind each skill work in new situations. The two depend on each other, yet i…

  132. Hugging Face Daily Papers TIER_1 English(EN) ·

    VLA强化学习的低秩结构

    Reinforcement learning (RL) is increasingly used to post-train vision-language-action (VLA) models, yet how RL reshapes these policies remains poorly understood. We find that RL across widely used flow-based VLA models, including π_{0.5} and GR00T~N1.5/N1.6, on LIBERO, ManiSkill,…

  133. Hugging Face Daily Papers TIER_1 English(EN) ·

    无标签引导:将测试时强化学习压缩到仅偏置子空间

    Test-time reinforcement learning (TTRL) enables models to improve their reasoning without relying on labeled training data, but existing approaches typically optimize a large fraction of the model parameters. This raises a natural question: can effective test-time adaptation emer…

  134. Hugging Face Daily Papers TIER_1 English(EN) ·

    在统一模型中通过交错强化学习学习原生反射

    Unified multimodal models can both look at and render images, so in principle they can repair their own generations: diagnose what an image gets wrong, revise it, observe the result, and diagnose again. Whether a revision helps is known only after it is rendered, so the reflectio…

  135. Hugging Face Daily Papers TIER_1 English(EN) ·

    QwenGyre:一种用于训练 xLong-Horizon Agent 的弹性强化学习框架

    Large language model (LLM) agents increasingly undertake extreme-long (xlong) horizon tasks, where a single execution can span hours, hundreds of model--environment interactions, and nearly 1M tokens per rollout. Applying online reinforcement learning (RL) to such executions pose…

  136. Hugging Face Daily Papers TIER_1 English(EN) ·

    玩转对等:用于可证明最优四边形块分解的强化学习

    A quadrilateral block decomposition of a planar domain is judged by whether it is complete, whether its elements are well shaped, and how many of its vertices are irregular. The last has a provable floor: the discrete Gauss-Bonnet identity enforces a lower bound on the total vert…

  137. arXiv cs.AI TIER_1 English(EN) · James Wu, Chris R. Sims ·

    强化学习中的策略复杂性、反应时间和有界理性

    arXiv:2609.28737v1 Announce Type: cross Abstract: Biological agents do not learn under conditions of unlimited computation. For humans, learning and choice are shaped by constraints on perception, attention, and working memory, which limit how much state information guides behavi…

  138. arXiv stat.ML TIER_1 English(EN) · Akira Kitaoka ·

    超越逆向优化:学习超越代理决策的目标函数

    arXiv:2610.09890v1 Announce Type: cross Abstract: Inverse optimization estimates the weights of an objective function that explain observed decisions as optimal solutions, and is used in a variety of fields. For mixed-integer linear programs (MILPs), existing methods aim to repro…

  139. arXiv cs.CV TIER_1 English(EN) · Bingyang Cui, Yujie Zhang, Yiling Xu, Yunfeng Guan ·

    OuroReward:文本到3D生成中强化学习的顺序奖励调度

    arXiv:2610.03423v1 Announce Type: new Abstract: Reinforcement learning (RL) for Text-to-3D (T23D) generation requires optimization across multiple quality dimensions such as semantic alignment and texture clarity. Existing methods typically optimize these dimensions simultaneousl…

  140. arXiv stat.ML TIER_1 English(EN) · Allen Tran, Jia Wan, Nathan Kallus, Aur\'elien Bibaut ·

    可扩展多任务逆强化学习

    arXiv:2610.00758v1 Announce Type: cross Abstract: By learning transferable rewards, inverse reinforcement learning (IRL) enables counterfactual evaluation of agents under modified environments. Such transfer places strict requirements on coverage since target environments affect …

  141. arXiv cs.CV TIER_1 English(EN) · Yifan Wang, Gordon Guocheng Qian, Yanyu Li, Anil Kag, Yun Fu ·

    Token-Level Video Reinforcement Learning

    arXiv:2610.01973v1 Announce Type: new Abstract: Reinforcement learning (RL) for video generation usually assigns one scalar reward to an entire sampled video. Yet a video is not uniformly flawed: some visual tokens may already satisfy the prompt, whereas others require correction…

  142. arXiv stat.ML TIER_1 English(EN) · Shashank Gupta, Pilhwa Lee ·

    Meta-reinforcement learning with minimum attention

    arXiv:2505.16741v5 Announce Type: replace-cross Abstract: Minimum attention applies the least action principle in changes of control concerning state and time, first proposed by Brockett. The involved regularization is highly relevant in emulating biological control, such as moto…

  143. arXiv stat.ML TIER_1 English(EN) · Toru Kitagawa, Shosei Sakaguchi, Aleksey Tetenov ·

    约束分类与策略学习

    arXiv:2106.12886v3 Announce Type: replace-cross Abstract: Modern machine learning approaches to classification, including AdaBoost, support vector machines, and deep neural networks, utilize surrogate loss techniques to circumvent the computational complexity of minimizing empiri…

  144. arXiv stat.ML TIER_1 English(EN) · Deqian Kong, Guangyan Sun, Sheng Cheng, Sirui Xie, Bo Pang, Jianwen Xie, Tony Geng, Caiwen Ding, Ying Nian Wu ·

    从随机探索中学习规划

    arXiv:2609.38383v1 Announce Type: cross Abstract: Random exploration reveals how an environment can be traversed before a goal is specified. Can this experience support long-range planning without policy-improvement training? Our random-walk analysis explains what temporal relati…

  145. arXiv stat.ML TIER_1 English(EN) · Dylan J. Foster, Alexander Rakhlin ·

    强化学习与交互式决策基础

    arXiv:2312.16730v2 Announce Type: replace-cross Abstract: Interactive decision making is the problem of learning to act well in an unknown environment, using the data that one's own actions generate to continuously improve, and arises in situations ranging from online platforms a…

  146. Towards AI TIER_1 English(EN) · Mayurnayak ·

    从零开始的强化学习 — 第一部分:导论

    <figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*bzw_rrScey_iRtRBlLodOg.png" /></figure><p>This is the Part-1 of Series <strong>Reinforcement Learning from Scratch. </strong>In this series we plan to cover the core concepts of modern RL from Q-table, Deep-Q Lea…

  147. Medium — Claude tag TIER_1 English(EN) · Arjun Keerthi ·

    Agentic Finetuning with Reinforcement Learning

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@arjunkeerthi1529.edu/agentic-finetuning-with-reinforcement-learning-b42e68857fbb?source=rss------claude-5"><img src="https://cdn-images-1.medium.com/max/1200/0*lNH-_f93luam3niK" width="1200" /…

  148. dev.to — LLM tag TIER_1 English(EN) · Prabhakar Chaudhary ·

    LoGRA:低秩梯度草图如何让LLM强化学习适配真实硬件

    <h1> LoGRA: How Low-Rank Gradient Sketches Make LLM Reinforcement Learning Fit on Real Hardware </h1> <p>Reinforcement learning post-training has become one of the most effective ways to improve large language model reasoning. Models like DeepSeek-R1 demonstrated that RL-based fi…

  149. dev.to — LLM tag TIER_1 English(EN) · Seth Wheeler ·

    将强化学习分类法与我自己的工具进行比对

    <p>Somebody handed me a list of reinforcement learning algorithms and asked whether any of them could help my tools. It is a ChatGPT session printed to PDF: roughly ninety algorithm names across fifteen sections, with no descriptions, no citations and no results. I want to be pre…