PulseAugur
实时 11:32:04

新研究探索用于智能体生存、导航和可解释性的高级强化学习 · 跟踪 7 个来源

研究人员正在探索强化学习(RL)的高级技术,以增强智能体的性能和可解释性。一项研究提出了程序化策略(PERL)作为进化 RL 中神经策略(NERL)的替代方案,证明了 PERL 更高的生存率。另一篇论文侧重于使用 RL 进行自主机器人的流感知导航,发现使用速度和记忆传感器训练的智能体优于那些具有显式全局流参数的智能体。此外,正在开发新的方法来解释 RL 智能体,例如使用归纳逻辑编程(ILP)为策略可解释性创建客观指标,以及一种称为 SteinGate 的技术,通过检测 RL 中罕见的灾难性尾部事件来提高安全性。 AI

影响 这些 RL 技术方面的进步可能带来更强大、更高效、更可解释的 AI 智能体,应用于从机器人技术到复杂决策场景的各种领域。

排序理由 多篇 arXiv 论文发表了关于强化学习内不同研究主题的论文。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 304 个来源。 我们如何撰写摘要 →

新研究探索用于智能体生存、导航和可解释性的高级强化学习 · 跟踪 7 个来源

报道来源 [304]

  1. arXiv cs.AI TIER_1 English(EN) · Yuxuan Zhu, Rohan Alur, Daniel Kang ·

    具有可验证奖励的强化学习的非平凡泛化界限

    arXiv:2607.14506v1 Announce Type: cross Abstract: While reinforcement learning with verifiable rewards (RLVR) is widely used to improve the reasoning capabilities of large language models (LLMs), the generalizability of the resulting models remains poorly understood. In this work…

  2. arXiv cs.LG TIER_1 English(EN) · Asha Ramanujam (Davidson School of Chemical Engineering, Purdue University, West Lafayette, IN), Adam Elyoumi (Davidson School of Chemical Engineering, Purdue University, West Lafayette, IN), Hao Chen (Davidson School of Chemical Engineering, Purdue Univ… ·

    SafeOR-Gym:面向实际运筹学问题的安全强化学习算法基准套件

    arXiv:2506.02255v2 Announce Type: replace Abstract: Most existing safe reinforcement learning (RL) benchmarks focus on robotics and control tasks, offering limited relevance to high-stakes domains that involve structured constraints, mixed-integer decisions, and industrial comple…

  3. arXiv cs.LG TIER_1 English(EN) · Zikun Zhang, Jiayuan Sheng, David D. Yao, Wenpin Tang ·

    用于微调离散扩散模型的连续时间强化学习框架

    arXiv:2607.14522v1 Announce Type: new Abstract: We formulate reinforcement learning (RL) in continuous time with discrete state spaces and possibly arbitrary action spaces via a stochastic control approach, where the state dynamics are modeled as a controlled continuous-time Mark…

  4. arXiv cs.CL TIER_1 English(EN) · Bowei He, Yankai Chen, Xiaokun Zhang, Xue Liu ·

    分支策略优化:沙盒原生语言代理强化学习

    arXiv:2607.14171v1 Announce Type: cross Abstract: Reinforcement learning has emerged as the dominant paradigm for training large language model (LLM) agents that interact with executable sandboxes. State-of-the-art algorithms such as PPO, RLOO, and GRPO inherit their rollout topo…

  5. arXiv cs.CL TIER_1 English(EN) · Jinyang Wu, Shuo Yang, Zhengxi Lu, Fan Zhang, Yuhao Shen, Lang Feng, Haoran Luo, Zheng Lian, Shuai Zhang, Zhengqi Wen, Jianhua Tao ·

    SEED: 智能体强化学习的自演进策略蒸馏

    arXiv:2607.14777v1 Announce Type: new Abstract: Large language models are increasingly trained as interactive agents for long-horizon tasks involving multi-turn interaction, tool use, and environment feedback. Outcome-based reinforcement learning (RL) provides a practical optimiz…

  6. arXiv cs.AI TIER_1 English(EN) · Mattie Fellows, Clarisse Wibault, Uljad Berdica, Johannes Forkel, Maike Osborne, Jakob N. Foerster ·

    完全离线强化学习

    arXiv:2505.22442v3 Announce Type: replace-cross Abstract: Offline RL (ORL) promises safe and sample-efficient deployment but existing methods rely on undocumented online interactions for hyperparameter tuning and lack reliable fully offline estimates of initial online performance…

  7. arXiv cs.AI TIER_1 English(EN) · Mohsen Amiri, Sindri Magn\'usson ·

    强化学习在切换非平稳马尔可夫决策过程中的应用:算法与收敛性分析

    arXiv:2503.18607v2 Announce Type: replace-cross Abstract: We introduce the Switching Non-Stationary Markov Decision Process (SNS-MDP) framework, in which the environment transitions among a finite set of MDPs governed by a latent Markov chain while the agent observes only the ext…

  8. arXiv cs.CL TIER_1 English(EN) · Jianhua Tao ·

    SEED: 智能体强化学习的自演进策略蒸馏

    Large language models are increasingly trained as interactive agents for long-horizon tasks involving multi-turn interaction, tool use, and environment feedback. Outcome-based reinforcement learning (RL) provides a practical optimization paradigm, but its sparse trajectory-level …

  9. arXiv cs.LG TIER_1 English(EN) · Junyi Wu, Dan Li ·

    Factorized Spectral Representations for Reinforcement Learning

    arXiv:2607.13498v1 Announce Type: new Abstract: Learning a compact model of the world from interaction data is central to sample-efficient deep reinforcement learning. Spectral representation methods have become the leading paradigm for representation learning in continuous contr…

  10. arXiv cs.AI TIER_1 English(EN) · Celeste Veronese, Edoardo Zorzi, Daniele Meli, Alessandro Farinelli ·

    通过归纳逻辑编程解释强化学习智能体

    arXiv:2607.13655v1 Announce Type: new Abstract: Explainable Reinforcement Learning (XRL) seeks to make Reinforcement Learning (RL) policies more transparent and interpretable, a key requirement in safety-critical and human-centric scenarios. However, it is mostly based on user st…

  11. arXiv cs.LG TIER_1 English(EN) · Andrea Maria Braghin, Nicol\`o Botteghi, Matteo Tomasetto, Andrea Manzoni, Gabriele Cazzulani ·

    基于强化学习的非定常流中的流感知最优导航

    arXiv:2607.13553v1 Announce Type: cross Abstract: Autonomous robotic navigation in nonstationary time-varying fluid flows remains a fundamental challenge due to partial observability and the unpredictability of realistic environments. While classical optimal control frameworks em…

  12. arXiv cs.LG TIER_1 English(EN) · Merlin Paul, Anup Aprem ·

    面向贝叶斯说服的结构化强化学习:在智能交互式驾驶中的应用

    arXiv:2607.13576v1 Announce Type: new Abstract: Interactive driving, wherein an intelligent lead vehicle equipped with real-time traffic data coordinates route choices of connected vehicles, offers a promising approach to dynamic traffic management. To address the challenge of ha…

  13. arXiv cs.AI TIER_1 Deutsch(DE) · Yassine Chemingui, Chenhua Fan, Honghao Wei, Janardhan Rao Doppa ·

    SteinGate:通过 Stein 散度实现尾部敏感的安全强化学习

    arXiv:2607.13175v1 Announce Type: cross Abstract: Safe reinforcement learning typically enforces safety by bounding expected cumulative costs, a criterion that often fails to detect rare but catastrophic tail events. To overcome these limitations, this paper introduces SteinGate,…

  14. arXiv cs.LG TIER_1 English(EN) · Jiale Fan, Andrei Cramariuc, Tifanny Portela, Marco Hutter ·

    用于运动的Actor-Critic强化学习中的预训练

    arXiv:2510.12363v4 Announce Type: replace-cross Abstract: The pretraining-finetuning paradigm has facilitated numerous transformative advancements in artificial intelligence research in recent years. However, in the domain of reinforcement learning (RL) for robot locomotion, indi…

  15. arXiv cs.LG TIER_1 English(EN) · Anton Roupassov-Ruiz, Yiyang Zuo ·

    进化强化学习中神经策略与程序策略的生存动力学

    arXiv:2601.04365v2 Announce Type: replace Abstract: In evolutionary reinforcement learning tasks (ERL), agent policies are often encoded as small artificial neural networks (NERL). Such representations lack explicit modular structure, limiting behavioral interpretation. We invest…

  16. arXiv cs.LG TIER_1 English(EN) · Wenpin Tang ·

    用于微调离散扩散模型的连续时间强化学习框架

    We formulate reinforcement learning (RL) in continuous time with discrete state spaces and possibly arbitrary action spaces via a stochastic control approach, where the state dynamics are modeled as a controlled continuous-time Markov chain (CTMC). We consider policy optimization…

  17. arXiv cs.LG TIER_1 English(EN) · Daniel Kang ·

    具有可验证奖励的强化学习的非平凡泛化界限

    While reinforcement learning with verifiable rewards (RLVR) is widely used to improve the reasoning capabilities of large language models (LLMs), the generalizability of the resulting models remains poorly understood. In this work, we establish the first non-vacuous generalizatio…

  18. Hugging Face Daily Papers TIER_1 English(EN) ·

    SEED: 智能体强化学习的自演化策略蒸馏

    Large language models are increasingly trained as interactive agents for long-horizon tasks involving multi-turn interaction, tool use, and environment feedback. Outcome-based reinforcement learning (RL) provides a practical optimization paradigm, but its sparse trajectory-level …

  19. arXiv cs.AI TIER_1 English(EN) · Alessandro Farinelli ·

    通过归纳逻辑编程解释强化学习智能体

    Explainable Reinforcement Learning (XRL) seeks to make Reinforcement Learning (RL) policies more transparent and interpretable, a key requirement in safety-critical and human-centric scenarios. However, it is mostly based on user studies, thus targeting the needs of a specific au…

  20. arXiv cs.CL TIER_1 English(EN) · Xue Liu ·

    分支策略优化:沙盒原生语言代理强化学习

    Reinforcement learning has emerged as the dominant paradigm for training large language model (LLM) agents that interact with executable sandboxes. State-of-the-art algorithms such as PPO, RLOO, and GRPO inherit their rollout topology from RLHF: for each prompt, N independent tra…

  21. arXiv cs.LG TIER_1 English(EN) · Anup Aprem ·

    面向贝叶斯说服的结构化强化学习:在智能交互式驾驶中的应用

    Interactive driving, wherein an intelligent lead vehicle equipped with real-time traffic data coordinates route choices of connected vehicles, offers a promising approach to dynamic traffic management. To address the challenge of harmonising decisions, this paper considers the st…

  22. arXiv cs.LG TIER_1 English(EN) · Gabriele Cazzulani ·

    基于强化学习的非定常流中的流感知最优导航

    Autonomous robotic navigation in nonstationary time-varying fluid flows remains a fundamental challenge due to partial observability and the unpredictability of realistic environments. While classical optimal control frameworks employed in robotics require unrealistic a-priori gl…

  23. arXiv cs.LG TIER_1 English(EN) · Dan Li ·

    Factorized Spectral Representations for Reinforcement Learning

    Learning a compact model of the world from interaction data is central to sample-efficient deep reinforcement learning. Spectral representation methods have become the leading paradigm for representation learning in continuous control by taking a matrix view of the transition ker…

  24. arXiv cs.AI TIER_1 English(EN) · A Run, Ziluo Ding ·

    非平稳性下的上下文强化学习:综述

    arXiv:2607.11906v1 Announce Type: new Abstract: The development of decision-pretrained transformers, algorithm distillation, long-context meta-RL, and retrieval-augmented agents has renewed interest in in-context reinforcement learning (ICRL): the ability of a pretrained or fine-…

  25. arXiv cs.LG TIER_1 English(EN) · Paolo Magliano, Puze Liu, Jan Peters, Davide Tateo, Raffaello Camoriano ·

    Safe Reinforcement Learning 中高效探索的方向性约束

    arXiv:2607.12784v1 Announce Type: cross Abstract: Reinforcement Learning has revolutionized the landscape of robotic research, allowing robust learning of complex robotic skills in simulation. However, real-world deployment in open-ended environments requires strong safety guaran…

  26. arXiv cs.LG TIER_1 English(EN) · Amber Srivastava ·

    强化学习中策略-环境协同设计的环境参数梯度定理

    arXiv:2607.12590v1 Announce Type: cross Abstract: Reinforcement learning (RL) is traditionally concerned with learning a control policy for a fixed environment. In many engineering systems, however, the environment itself is alterable: physical or operational parameters can be tu…

  27. arXiv cs.AI TIER_1 English(EN) · Zhouchonghao Wu, Raymond Song, Vedant Mundheda, Luis E. Navarro-Serment, Christof Schoenborn, Jeff Schneider ·

    TADPO:强化学习走向越野

    arXiv:2603.05995v2 Announce Type: replace-cross Abstract: Off-road autonomous driving poses significant challenges such as navigating unmapped, variable terrain with uncertain and diverse dynamics. Addressing these challenges requires effective long-horizon planning and adaptable…

  28. arXiv cs.AI TIER_1 English(EN) · Tongxi Wang, Zhuoyang Xia, Xinran Chen, Shan Liu ·

    追踪漂移:非平稳强化学习的变异感知熵调度

    arXiv:2601.19624v3 Announce Type: replace-cross Abstract: Real-world reinforcement learning often faces environment drift, but most existing methods rely on static entropy coefficients/target entropy, causing over-exploration during stable periods and under-exploration after drif…

  29. arXiv cs.AI TIER_1 English(EN) · Xin Qiu, Yulu Gan, Conor F. Hayes, Qiyao Liang, Yinggan Xu, Roberto Dailey, Elliot Meyerson, Babak Hodjat, Risto Miikkulainen ·

    大规模进化策略:超越强化学习的LLM微调

    arXiv:2509.24372v3 Announce Type: replace-cross Abstract: Fine-tuning large language models (LLMs) for downstream tasks is an essential stage of modern AI deployment. Reinforcement learning (RL) has emerged as the dominant fine-tuning paradigm, underpinning many state-of-the-art …

  30. arXiv cs.AI TIER_1 English(EN) · Emil Mittag, Richard Dazeley, Peter Vamplew ·

    OOD-RL-Bench:强化学习中分布外检测的基准框架

    arXiv:2607.12523v1 Announce Type: cross Abstract: Reliable reinforcement learning (RL) agents must maintain operational integrity amidst sensor malfunctions, dynamic disturbances, and slow environmental shifts. The detection of out-of-distribution conditions is pivotal to determi…

  31. arXiv cs.AI TIER_1 English(EN) · Jonas Ehrhardt, Ren\'e Heesch, Oliver Niggemann ·

    面向参数化动作马尔可夫决策过程的知识与梯度引导强化学习

    arXiv:2607.12924v1 Announce Type: new Abstract: In this paper, we study Reinforcement Learning in Parametrized Action Markov Decision Processes (PAMDP), where each decision consists of a symbolic action and numerical parameters. In such settings Reinforcement Learning algorithms …

  32. arXiv cs.AI TIER_1 English(EN) · Yuhui Bie, Guowei Xu, Yaojun Wang ·

    面向智能温室强化学习控制的先校准后奖励组件审计

    arXiv:2607.11959v1 Announce Type: new Abstract: Greenhouse reinforcement learning can test climate-control ideas at a speed and scale that is difficult to achieve with crop experiments alone. For smart-greenhouse control, however, a single simulator return is not enough: a grower…

  33. arXiv cs.AI TIER_1 English(EN) · Oliver Niggemann ·

    面向参数化动作马尔可夫决策过程的知识与梯度引导强化学习

    In this paper, we study Reinforcement Learning in Parametrized Action Markov Decision Processes (PAMDP), where each decision consists of a symbolic action and numerical parameters. In such settings Reinforcement Learning algorithms typically determine parameters with one-shot est…

  34. arXiv cs.LG TIER_1 English(EN) · Raffaello Camoriano ·

    安全强化学习中高效探索的方向性约束

    Reinforcement Learning has revolutionized the landscape of robotic research, allowing robust learning of complex robotic skills in simulation. However, real-world deployment in open-ended environments requires strong safety guarantees to prevent dangerous or harmful behaviors. Sa…

  35. Hugging Face Daily Papers TIER_1 English(EN) ·

    Safe Reinforcement Learning 中的高效探索定向约束

    Reinforcement Learning has revolutionized the landscape of robotic research, allowing robust learning of complex robotic skills in simulation. However, real-world deployment in open-ended environments requires strong safety guarantees to prevent dangerous or harmful behaviors. Sa…

  36. arXiv cs.LG TIER_1 English(EN) · Amber Srivastava ·

    Reinforcement Learning中策略-环境协同设计的环境参数梯度定理

    Reinforcement learning (RL) is traditionally concerned with learning a control policy for a fixed environment. In many engineering systems, however, the environment itself is alterable: physical or operational parameters can be tuned to shape the transition dynamics and costs exp…

  37. arXiv cs.AI TIER_1 English(EN) · Peter Vamplew ·

    OOD-RL-Bench:强化学习中分布外检测的基准框架

    Reliable reinforcement learning (RL) agents must maintain operational integrity amidst sensor malfunctions, dynamic disturbances, and slow environmental shifts. The detection of out-of-distribution conditions is pivotal to determining when an agent's observations, transitions, or…

  38. Hugging Face Daily Papers TIER_1 English(EN) ·

    OOD-RL-Bench:强化学习中分布外检测的基准框架

    Reliable reinforcement learning (RL) agents must maintain operational integrity amidst sensor malfunctions, dynamic disturbances, and slow environmental shifts. The detection of out-of-distribution conditions is pivotal to determining when an agent's observations, transitions, or…

  39. arXiv cs.LG TIER_1 English(EN) · Jiaheng Hu, Zizhao Wang, Peter Stone, Roberto Mart\'in-Mart\'in ·

    高效分层强化学习中的解耦无监督技能发现

    arXiv:2410.11251v2 Announce Type: replace Abstract: A hallmark of intelligent agents is the ability to learn reusable skills purely from unsupervised interaction with the environment. However, existing unsupervised skill discovery methods often learn entangled skills where one sk…

  40. arXiv cs.LG TIER_1 English(EN) · Ekkachai Jueng ·

    Q-Learning Lab:通过学习者生成的轨迹分析教授强化学习

    arXiv:2607.10802v1 Announce Type: cross Abstract: Reinforcement learning is usually introduced through the Bellman update, yet the equation often remains abstract to undergraduates: they watch policy arrows converge but rarely observe how each value is computed or why an action i…

  41. arXiv cs.LG TIER_1 English(EN) · Simone Drago, Marco Mussi, Leonardo Bianconi, Alberto Maria Metelli ·

    通用偏好基强化学习:不可比性的理性模型

    arXiv:2607.11432v1 Announce Type: new Abstract: In this work, we study the reinforcement learning (RL) problem from pairwise trajectory comparisons provided by a human expert. We generalize preference-based RL by formalizing a novel setting in which the expert can also label traj…

  42. arXiv cs.LG TIER_1 English(EN) · Sai Anirudh Katupilla, Shreeya Dasa Lakshminath ·

    基于数据并行Gibbs采样的非参数贝叶斯逆强化学习

    arXiv:2607.09886v1 Announce Type: new Abstract: Inverse Reinforcement Learning recovers reward functions from expert demonstrations, but standard formulations assume that all demonstrations come from a single expert. When demonstrations are pooled from multiple experts with disti…

  43. arXiv cs.AI TIER_1 English(EN) · Bonan Wang, Letian Tao, Bin Shuai, Jiaxin Gao, Wenxin Zhao, Wei Xiong, Kehua Sheng, Bo Zhang, Yang Guan, Shengbo Eben Li ·

    FAST:用于自动驾驶并行强化学习的对齐采样与训练框架

    arXiv:2606.21587v2 Announce Type: replace-cross Abstract: Deep reinforcement learning is pivotal for closed-loop autonomous driving yet remains constrained by severe bottlenecks in sampling efficiency. Standard parallel sampling mitigates this but suffers from the straggler effec…

  44. arXiv cs.AI TIER_1 English(EN) · Xiaoya Li, Albert Wang, Guoyin Wang, Chris Shum, Jiwei Li ·

    CRINN:用于近似最近邻搜索的对比强化学习

    arXiv:2508.02091v3 Announce Type: replace-cross Abstract: Approximate nearest-neighbor search (ANNS) algorithms have become increasingly critical for recent AI applications, particularly in retrieval-augmented generation (RAG) and agent-based LLM applications. In this paper, we p…

  45. arXiv cs.AI TIER_1 English(EN) · John Wikman, Alexandre Proutiere, David Broman ·

    不可观测随机延迟的自适应强化学习

    arXiv:2506.14411v2 Announce Type: replace-cross Abstract: In standard reinforcement learning (RL) settings, the interaction between the agent and the environment is typically modeled as a Markov decision process (MDP), which assumes that the agent observes the system state instan…

  46. arXiv cs.AI TIER_1 English(EN) · Pengfei Cai, Utkarsh Utkarsh, Alan Edelman, Christopher Vincent Rackauckas, Rafael Gomez-Bombarelli ·

    基于可验证物理学的强化学习:使用连续奖励进行LLM的训练后学习

    arXiv:2607.10474v1 Announce Type: cross Abstract: Partial differential equations (PDEs) are foundational to modeling in science and engineering, but constructing reliable numerical solvers remains labor-intensive, demanding expert knowledge of discretization schemes, stability co…

  47. arXiv cs.AI TIER_1 English(EN) · Wenke Xia, Pei Ren, Wenbo Yu, Yizhuo Zhang, Jifan Li, Yixue Zhang, Yinuo Zhao, Qingyang Gao, Jianlong Fu, Jian Tang, Ji-Rong Wen, Zhengping Che, Di Hu ·

    Robo-ValueRL:离线到在线强化学习的可靠价值估计

    arXiv:2607.09866v1 Announce Type: cross Abstract: Offline-to-online reinforcement learning is promising for generalizable robotic manipulation, yet its full-stack complexity obscures reproduction and diagnosis. Within such systems, value estimation plays a central role in priorit…

  48. arXiv cs.AI TIER_1 English(EN) · Roberto Garrone ·

    Adaptive Multi-Agent Systems in Policy Regimes Across Transfer Learning

    arXiv:2607.09685v1 Announce Type: cross Abstract: Policy models often assume that the relationship between a policy instrument and its outcome remains stable across institutional conditions. In adaptive socio-technical systems this assumption may fail: regulatory change can alter…

  49. arXiv cs.AI TIER_1 English(EN) · Mingyuan Wu, Jingcheng Yang, Shengyi Qian, Xudong Wang, Jize Jiang, Qifan Wang, Aashu Singh, Khoi Pham, Fei Liu, Zhaolun Su, Zhuokai Zhao, Klara Nahrstedt, Jianyu Wang, Hanchao Yu ·

    SVR-R1:在强化学习中通过自验证引导多模态推理

    arXiv:2607.10966v1 Announce Type: new Abstract: We introduce Self-Verified Reasoner (SVR-R1), a multi-turn RL framework that turns a model's own verification into a learning signal for multimodal reasoning. For each query, the model proposes an answer using the same weights, and …

  50. arXiv cs.AI TIER_1 English(EN) · Mianqiu Huang, Taofeng Xue, Chong Peng, Jinrui Ding, Sicheng Fan, Jiale Hong, Yufei Gao, Xiaocheng Zhang, Linsen Guo, Xin Yang, Dengchang Zhao, Yuchen Xie, Peng Pei, Xunliang Xie, Xipeng Qiu ·

    EvoCUA-1.5:面向多轮计算机使用代理的在线强化学习

    arXiv:2607.09773v1 Announce Type: new Abstract: Computer-use agents must solve long-horizon tasks through repeated interaction with partially observable, multimodal desktop environments. Although imitation learning and offline trajectory refinement provide strong priors, static t…

  51. arXiv cs.AI TIER_1 English(EN) · Rongping Zhou, Omid Tavallaie, Shuaijun Chen, Albert Y. Zomaya ·

    衡量模拟到真实差距:为 AIoT 系统中的强化学习设计经济实惠的真实世界基准测试平台

    arXiv:2607.10309v1 Announce Type: new Abstract: Reinforcement learning (RL) is commonly employed to enhance the performance of autonomous systems, including the Autonomous Internet of Things (AIoT). However, the trial-and-error nature of RL, when conducted in real-world environme…

  52. arXiv cs.AI TIER_1 English(EN) · Rushuai Yang, Zhuo Han, Houlin Li, Hecheng Wang, Zhichao Wu, Rui Zhang, Zhaowei Zhang, Zihong Chen, Xiaohan Yan, Chiming Liu, Yi Chen, Wei Shan, Maoqing Yao ·

    VINE:为强化学习驯服生成式控制策略

    arXiv:2607.10369v1 Announce Type: cross Abstract: Flow-matching policies have emerged as an effective policy parameterization for robot learning. They iteratively generate actions from noise, enabling highly expressive modeling of complex and multimodal action distributions. Howe…

  53. arXiv cs.AI TIER_1 English(EN) · Baohe Zhang, Lilli Frison, Thomas Brox, Joschka B\"odecker ·

    受限强化学习用于安全热泵控制

    arXiv:2409.19716v2 Announce Type: replace-cross Abstract: Constrained Reinforcement Learning (RL) has emerged as a significant research area within RL, where integrating constraints with rewards is crucial for enhancing safety and performance across diverse control tasks. In the …

  54. arXiv cs.AI TIER_1 English(EN) · DeepReinforce Team, Xiaoya Li, Guoyin Wang, Songqiao Su, Chris Shum, Jiwei Li ·

    GrandCode:通过代理强化学习在竞技编程中达到特级大师级别

    arXiv:2604.02721v2 Announce Type: replace Abstract: Competitive programming remains one of the last few human strongholds in coding against AI. The best AI system to date still underperforms the best humans competitive programming: the most recent best result, Google's Gemini~3 D…

  55. arXiv cs.AI TIER_1 English(EN) · Yunhai Feng, Natalie Leung, Jiaxuan Wang, Lujie Yang, Haozhi Qi, Preston Culbertson ·

    一种极简的重定向引导强化学习配方,用于灵巧操作

    arXiv:2607.11874v1 Announce Type: cross Abstract: Recent work in humanoid whole-body control has found success with a simple recipe: retarget human motion to robot kinematic references, then train policies via reinforcement learning (RL) to track them. But how does this recipe tr…

  56. arXiv cs.AI TIER_1 English(EN) · Alper Kamil Bozkurt, Shangtong Zhang, Yuichi Motai ·

    主动式离线到在线强化学习

    arXiv:2607.11720v1 Announce Type: cross Abstract: Background: Offline reinforcement learning (RL) enables effective policies to be trained from large, previously collected datasets and subsequently improved through limited online interaction. This offline-to-online RL (O2O-RL) pa…

  57. arXiv cs.AI TIER_1 English(EN) · Preston Culbertson ·

    一种极简的重定向引导强化学习配方,用于灵巧操作

    Recent work in humanoid whole-body control has found success with a simple recipe: retarget human motion to robot kinematic references, then train policies via reinforcement learning (RL) to track them. But how does this recipe transfer to dexterous manipulation? The answer is no…

  58. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Alfred Höß ·

    用于 C-V2X RAT 选择的多智能体强化学习

    Vehicles are increasingly equipped with advanced V2X communication capabilities. While early V2X apps utilized services such as Cooperative Awareness Messages, recent developments have allowed more advanced applications including cooperative driving, shared perception, and sensor…

  59. arXiv cs.AI TIER_1 English(EN) · Yuichi Motai ·

    主动式离线到在线强化学习

    Background: Offline reinforcement learning (RL) enables effective policies to be trained from large, previously collected datasets and subsequently improved through limited online interaction. This offline-to-online RL (O2O-RL) paradigm is particularly promising in nonstationary …

  60. Hugging Face Daily Papers TIER_1 English(EN) ·

    主动式离线到在线强化学习

    Background: Offline reinforcement learning (RL) enables effective policies to be trained from large, previously collected datasets and subsequently improved through limited online interaction. This offline-to-online RL (O2O-RL) paradigm is particularly promising in nonstationary …

  61. arXiv cs.LG TIER_1 English(EN) · Alberto Maria Metelli ·

    通用偏好基强化学习:不可比性的理性模型

    In this work, we study the reinforcement learning (RL) problem from pairwise trajectory comparisons provided by a human expert. We generalize preference-based RL by formalizing a novel setting in which the expert can also label trajectory pairs as incomparable, i.e., when neither…

  62. arXiv cs.AI TIER_1 English(EN) · Jiayu Yao, Yiwei Wang, Anmeng Zhang, Zhe Sun, Songsong Wang, Lingrui Mei, Yuyao Ge, Shenghua Liu ·

    强化学习中的多模态奖励黑客

    arXiv:2607.09492v1 Announce Type: new Abstract: Reinforcement learning (RL) is increasingly used to align multimodal large language models (MLLMs), but higher rewards do not always imply better task performance. This risk is amplified when visual evidence is evaluated by text-onl…

  63. arXiv cs.AI TIER_1 English(EN) · Guanquan Wang, Yoshimasa Tsuruoka ·

    高效离线强化学习的捷径轨迹规划

    arXiv:2607.09336v1 Announce Type: cross Abstract: Diffusion-based trajectory planners have shown strong performance in offline reinforcement learning, but their iterative denoising process often incurs high inference cost. Consistency-based planners reduce the number of sampling …

  64. arXiv cs.AI TIER_1 English(EN) · Brent Kong, Tejas Ram, Tony Yue Yu ·

    AlphaZero 在稀疏奖励游戏中的应用:局限性与辅助监督

    arXiv:2607.08984v1 Announce Type: cross Abstract: AlphaZero has demonstrated that a neural-guided Monte Carlo Tree Search can achieve superhuman performance, but strong play does not necessarily imply perfect play. We study this gap in two oracle-evaluable domains with contrastin…

  65. arXiv cs.LG TIER_1 English(EN) · Alan Nadelsticher Ruvalcaba ·

    深度强化学习中发育奖励机制的进化发现

    arXiv:2606.20858v2 Announce Type: replace Abstract: The temporal structure of reward composition in reinforcement learning (RL) is typically hand-designed and held fixed throughout training, leaving the progression of motivational priorities largely unexplored. In this work, we p…

  66. arXiv cs.LG TIER_1 English(EN) · Iris Xu, Sunshine Jiang, John Marangola, Nitish Dashora, Richard Li, Thomas Liu, Zexue He, Yuheng Zhi, Alex Pentland, Pulkit Agrawal, Zhang-Wei Hong ·

    从少中学习更多:事后强化学习

    arXiv:2607.09042v1 Announce Type: new Abstract: Reinforcement learning (RL) is increasingly used to post-train vision-language-action (VLA) models, but every update consumes robot rollouts that are slow and costly to collect, making sample efficiency a central concern. Manipulati…

  67. arXiv cs.LG TIER_1 English(EN) · Elham Daneshmand, Majid Khadiv, Glen Berseth, Hsiu-Chin Lin ·

    SafeExplorer:一种用于带恢复干预的强化学习的无偏策略梯度

    arXiv:2607.08925v1 Announce Type: new Abstract: Training reinforcement-learning agents directly on physical robots makes every fall costly, since a fall can damage the platform and cannot be undone like a simulator reset; the goal is therefore to minimize falls during training ra…

  68. arXiv cs.AI TIER_1 English(EN) · Tanmay Ambadkar, Sourav Panda, Shreyash Kale, Jonathan Dodge, Abhinav Verma ·

    Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization

    arXiv:2602.07764v2 Announce Type: replace-cross Abstract: Multi-objective reinforcement learning (MORL) seeks to train agents capable of balancing conflicting objectives. While single preference-conditioned policies offer a highly scalable solution, existing approaches remain bri…

  69. arXiv cs.AI TIER_1 English(EN) · Zihan Zhou, Animesh Garg, Ajay Mandlekar, Caelan Garrett ·

    ReinforceGen:混合技能策略,结合自动化数据生成与强化学习

    arXiv:2512.16861v2 Announce Type: replace-cross Abstract: Long-horizon manipulation has been a long-standing challenge in the robotics community. We propose ReinforceGen, a system that combines task decomposition, data generation, imitation learning, and motion planning to form a…

  70. arXiv cs.AI TIER_1 English(EN) · Shenghua Liu ·

    强化学习中的多模态奖励黑客

    Reinforcement learning (RL) is increasingly used to align multimodal large language models (MLLMs), but higher rewards do not always imply better task performance. This risk is amplified when visual evidence is evaluated by text-only or weakly grounded rewards. We study reward ha…

  71. arXiv cs.AI TIER_1 English(EN) · Yoshimasa Tsuruoka ·

    高效离线强化学习的捷径轨迹规划

    Diffusion-based trajectory planners have shown strong performance in offline reinforcement learning, but their iterative denoising process often incurs high inference cost. Consistency-based planners reduce the number of sampling steps, yet they typically rely on a two-stage teac…

  72. arXiv cs.AI TIER_1 English(EN) · Nivasini Ananthakrishnan, Mark Bedaywi, Michael I. Jordan, Stuart Russell, Nika Haghtalab ·

    Provably Optimal Learning Algorithms for Assistance Games

    arXiv:2607.08012v1 Announce Type: cross Abstract: This paper studies an online variant of the assistance games framework, where an informed agent and an uninformed agent repeatedly interact over $T$ timesteps to optimize a common reward function. While the informed agent (the hum…

  73. arXiv cs.AI TIER_1 English(EN) · Ezgi Korkmaz ·

    深度强化学习评估与设计范式的原则性分析

    arXiv:2607.07769v1 Announce Type: cross Abstract: Starting from the utilization of deep neural networks to approximate the state-action value function that led to winning one of the most challenging games, to algorithmic advancements that allowed solving problems without even exp…

  74. arXiv cs.AI TIER_1 English(EN) · Mumuksh Tayal, Manan Tayal, Ravi Prakash ·

    Safe Flow Q-Learning:基于可达性的流策略的离线安全强化学习

    arXiv:2603.15136v2 Announce Type: replace-cross Abstract: Offline safe reinforcement learning (RL) seeks reward-maximizing policies from static datasets under strict safety constraints. Existing methods often rely on soft expected-cost objectives or iterative generative inference…

  75. arXiv cs.AI TIER_1 English(EN) · Xiuyi Lou, Zicheng Xu, Yu-Neng Chuang, Hoang Anh Duy Le, Zhaozhuo Xu, Guanchu Wang, Vladimir Braverman ·

    当不切实际的 Token 被强化:LLM 强化学习的尾部感知信用校准

    arXiv:2607.07976v1 Announce Type: cross Abstract: Reinforcement learning (RL) has achieved remarkable success in enhancing the reasoning capabilities of large language models (LLMs). However, widely used critic-free RL methods rely on uniform credit assignment, broadcasting the s…

  76. arXiv cs.LG TIER_1 English(EN) · Nika Haghtalab, Chara Podimata, Kunhe Yang ·

    校准的Stackelberg博弈:学习对抗校准智能体的最优承诺

    arXiv:2306.02704v2 Announce Type: replace-cross Abstract: We introduce \emph{Calibrated Stackelberg Games (CSGs)}, a generalization of the standard Stackelberg Games (SGs) framework. In CSGs, a principal repeatedly interacts with an agent who (contrary to standard SGs) does not h…

  77. arXiv cs.CL TIER_1 English(EN) · Weiqi Wang, Xin Liu, Binxuan Huang, Hejie Cui, Rongzhi Zhang, Changlong Yu, Shuowei Jin, Jingfeng Yang, Qingyu Yin, Zhengyang Wang, Zheng Li, Yifan Gao, Priyanka Nigam, Bing Yin, Lihong Li, Yangqiu Song ·

    HeaPA:面向LLM强化学习的感知难度堆栈采样和在线查询增强

    arXiv:2601.22448v2 Announce Type: replace-cross Abstract: RLVR has become a standard recipe for training LLMs on reasoning tasks with verifiable outcomes, but when rollout generation dominates the cost, efficiency hinges on which prompts are sampled and when. In practice, prompt …

  78. arXiv cs.LG TIER_1 English(EN) · Zhang-Wei Hong ·

    从少中学习更多:事后强化学习

    Reinforcement learning (RL) is increasingly used to post-train vision-language-action (VLA) models, but every update consumes robot rollouts that are slow and costly to collect, making sample efficiency a central concern. Manipulation tasks typically provide only sparse rewards, …

  79. arXiv cs.LG TIER_1 English(EN) · Tony Yue Yu ·

    AlphaZero 在稀疏奖励游戏中:局限性与辅助监督

    AlphaZero has demonstrated that a neural-guided Monte Carlo Tree Search can achieve superhuman performance, but strong play does not necessarily imply perfect play. We study this gap in two oracle-evaluable domains with contrasting structure: Connect Four, a solved partisan game …

  80. arXiv cs.AI TIER_1 English(EN) · Zhenyu Hou, Yujiang Li, Jie Tang, Yuxiao Dong ·

    面向Agentic强化学习的单次部署异步优化

    arXiv:2607.07508v1 Announce Type: cross Abstract: Reinforcement learning (RL) is becoming increasingly important for post-training large language models (LLMs). Previous RL pipelines for LLMs were mostly synchronous and batch-interleaved, which is inefficient for long-horizon age…

  81. arXiv cs.LG TIER_1 English(EN) · Hokyun Im, Andrey Kolobov, Jianlong Fu, Youngwoon Lee ·

    通过一步流策略进行潜在策略引导

    arXiv:2603.05296v2 Announce Type: replace-cross Abstract: Offline reinforcement learning (RL) allows robots to learn from offline datasets without risky exploration. Yet, offline RL's performance often hinges on a brittle trade-off between (1) return maximization, which can push …

  82. arXiv cs.LG TIER_1 English(EN) · Denis Belomestny, Alexander Gasnikov, Egor Gladin, Alexey Naumov, Artemy Rubtsov, Yuri Sapronov, Daniil Tiapkin, Nikita Yudin ·

    强化学习的数学方法

    arXiv:2607.06935v1 Announce Type: cross Abstract: Reinforcement learning (RL) is increasingly grounded in tools from probability, optimization, and operator theory. This survey organizes the mathematical structures that underpin the design and analysis of modern algorithms in RL.…

  83. arXiv cs.LG TIER_1 English(EN) · Georg Sch\"afer, Jakob Rehrl, Stefan Huber, Simon Hirlaender ·

    使用模型预测控制的思想实现安全强化学习

    arXiv:2607.07252v1 Announce Type: new Abstract: Reinforcement learning (RL) enables the synthesis of control policies directly from data, making it highly appealing for complex cyber-physical systems (CPSs) and robotics. A persistent challenge, however, is ensuring strict, hard s…

  84. arXiv cs.AI TIER_1 English(EN) · Ignacio D. Lopez-Miguel, Ezio Bartocci, Thomas Eiter, Martin Tappler ·

    ORCAID:深度强化学习策略的斜率规则连续动作解释

    arXiv:2607.07235v1 Announce Type: cross Abstract: Explainability remains a key issue in reinforcement learning (RL). Distilling an interpretable policy from an agent trained in a complex environment is particularly challenging when the action space is continuous. We introduce ORC…

  85. arXiv cs.AI TIER_1 English(EN) · Zetian Hu, Shunyu Liu, Junjie Zhang, Yongcheng Jing, Ting-En Lin, Yongbin Li, Dacheng Tao ·

    多任务代理强化学习的熵节奏策略优化

    arXiv:2607.07178v1 Announce Type: cross Abstract: Recent breakthroughs of Reinforcement Learning (RL) have highlighted its potential for complex agentic Large Language Model (LLM) tasks. However, existing efforts largely focus on single-task settings, whereas real-world deploymen…

  86. arXiv cs.AI TIER_1 English(EN) · Dennis Gross, Quentin Mazouni, Helge Spieker, Arnaud Gotlieb ·

    Gimitest:用于测试强化学习策略的综合工具

    arXiv:2607.07029v1 Announce Type: cross Abstract: Reinforcement learning (RL) policies can be unsafe and vulnerable to attacks. Ensuring their reliability is often a pain point as existing automated testing methods target only selected environments, testing scenarios, and RL algo…

  87. arXiv cs.CL TIER_1 English(EN) · Vladimir Braverman ·

    当不可思议的 Token 被强化:LLM 强化学习的尾部感知信用校准

    Reinforcement learning (RL) has achieved remarkable success in enhancing the reasoning capabilities of large language models (LLMs). However, widely used critic-free RL methods rely on uniform credit assignment, broadcasting the same advantage to all tokens regardless of their di…

  88. Hugging Face Daily Papers TIER_1 English(EN) ·

    Agentic Reinforcement Learning 的单次部署异步优化

    Reinforcement learning (RL) is becoming increasingly important for post-training large language models (LLMs). Previous RL pipelines for LLMs were mostly synchronous and batch-interleaved, which is inefficient for long-horizon agentic tasks. Recently, asynchronous RL has emerged …

  89. arXiv cs.AI TIER_1 English(EN) · Yuxiao Dong ·

    面向Agentic强化学习的单回滚异步优化

    Reinforcement learning (RL) is becoming increasingly important for post-training large language models (LLMs). Previous RL pipelines for LLMs were mostly synchronous and batch-interleaved, which is inefficient for long-horizon agentic tasks. Recently, asynchronous RL has emerged …

  90. Hugging Face Daily Papers TIER_1 English(EN) ·

    使用模型预测控制的思路实现安全强化学习

    Reinforcement learning (RL) enables the synthesis of control policies directly from data, making it highly appealing for complex cyber-physical systems (CPSs) and robotics. A persistent challenge, however, is ensuring strict, hard safety constraints during the active learning pha…

  91. arXiv cs.LG TIER_1 English(EN) · Simon Hirlaender ·

    使用模型预测控制的思想实现安全强化学习

    Reinforcement learning (RL) enables the synthesis of control policies directly from data, making it highly appealing for complex cyber-physical systems (CPSs) and robotics. A persistent challenge, however, is ensuring strict, hard safety constraints during the active learning pha…

  92. arXiv cs.AI TIER_1 English(EN) · Martin Tappler ·

    ORCAID:深度强化学习策略的斜率规则连续动作解释

    Explainability remains a key issue in reinforcement learning (RL). Distilling an interpretable policy from an agent trained in a complex environment is particularly challenging when the action space is continuous. We introduce ORCAID, a novel method for extracting interpretable r…

  93. Hugging Face Daily Papers TIER_1 English(EN) ·

    多任务代理强化学习的熵节奏策略优化

    Recent breakthroughs of Reinforcement Learning (RL) have highlighted its potential for complex agentic Large Language Model (LLM) tasks. However, existing efforts largely focus on single-task settings, whereas real-world deployment necessitates a generalist agent capable of solvi…

  94. arXiv cs.AI TIER_1 English(EN) · Dacheng Tao ·

    多任务代理强化学习的熵节奏策略优化

    Recent breakthroughs of Reinforcement Learning (RL) have highlighted its potential for complex agentic Large Language Model (LLM) tasks. However, existing efforts largely focus on single-task settings, whereas real-world deployment necessitates a generalist agent capable of solvi…

  95. arXiv cs.AI TIER_1 English(EN) · Arnaud Gotlieb ·

    Gimitest:用于测试强化学习策略的综合性工具

    Reinforcement learning (RL) policies can be unsafe and vulnerable to attacks. Ensuring their reliability is often a pain point as existing automated testing methods target only selected environments, testing scenarios, and RL algorithms. To address this, we propose a comprehensiv…

  96. arXiv cs.AI TIER_1 English(EN) · Muhammad Zain Amin, Kibele Sebnem Yildirim ·

    具有跨回合记忆和策略蒸馏的自我评估强化学习 (SRRL)

    arXiv:2607.05541v1 Announce Type: cross Abstract: Reinforcement Learning is commonly used to train large language models using environmental feedback. In applied settings, the environment usually provides sparse or delayed feedback. This makes it difficult for the model to pinpoi…

  97. arXiv cs.AI TIER_1 English(EN) · Yuhang Zhou, Kai Zheng, Haoling Li, Dengyun Peng, Can Xu, Jingjing Chen ·

    TurnOPD:使on-policy蒸馏具备回合感知能力,以实现高效的长视域智能体训练

    arXiv:2607.05804v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student policy by matching a stronger teacher on the student's own trajectories, offering a promising framework for language agent training. However, its application to long-horizon agentic task…

  98. arXiv cs.AI TIER_1 English(EN) · Akshay Arora, Ishan Nigam, Ashutosh Aggarwal, Shefali Bansal, Krishna Singh, Sweta Kumari, Nikhil Mittal, Shariq Farhan, Siddarth Malreddy ·

    超越静态评估:为可扩展的Agentic强化学习构建模拟环境

    arXiv:2607.05773v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve into autonomous agents, traditional static evaluation fails to capture multi-step decision-making. We introduce AgenticAI-Supervisor, an API and UI-driven RL Gym environment that decouples envi…

  99. arXiv cs.AI TIER_1 English(EN) · Haiwen Yi, Xinyuan Song ·

    利用离线强化学习控制LLM Agent的工具

    arXiv:2607.05458v1 Announce Type: cross Abstract: Large language model (LLM) agents are usually improved by changing prompts, models, or hand-written workflows, while the execution harness around the model is treated as fixed infrastructure. We argue that this harness is itself a…

  100. arXiv cs.LG TIER_1 English(EN) · Gojko Perovic, Nuno Ferreira Duarte, Atabak Dehban, Gon\c{c}alo Teixeira, Egidio Falotico, Jos\'e Santos-Victor ·

    HERB:面向装箱问题的增强学习高效人类增强强化学习

    arXiv:2504.16595v2 Announce Type: replace-cross Abstract: Packing objects efficiently is a fundamental problem in logistics, warehouse automation, and robotics. When dealing with highly diverse 3D objects (household or grocery items), closed-form solutions are infeasible, and heu…

  101. arXiv cs.LG TIER_1 English(EN) · Shuze Daniel Liu, Claire Chen, Jiabao Sean Xiao, Xin Chen, David Simchi-Levi ·

    多买方市场中的战略议价:基于可验证奖励的强化学习用于LLM谈判

    arXiv:2607.05863v1 Announce Type: new Abstract: Negotiation is a fundamental strategic interaction in management science, characterized by agents attempting to reach agreements while protecting private information, such as reservation costs and hidden valuations. A prevalent yet …

  102. arXiv cs.AI TIER_1 English(EN) · Alexander Rombach, Chantale Lauer, Nijat Mehdiyev ·

    通过强化学习提升LLM生成过程模型质量:奖励函数设计的作用

    arXiv:2607.06175v1 Announce Type: cross Abstract: Large language models (LLMs) can generate BPMN process models from natural-language descriptions, yet supervised fine-tuning (SFT) limits their output quality to the patterns present in the training data. Reinforcement learning (R…

  103. arXiv cs.AI TIER_1 English(EN) · Yijun Zhang, Fan Xu, Jiaxin Ding, Yule Xie, Shiqing Gao, Xin Ding, Haoxiang Zhang, Luoyi Fu, Xinbing Wang ·

    基于信息增益的部署策略优化:一种用于多轮LLM智能体的自适应树状部署方法

    arXiv:2607.06223v1 Announce Type: new Abstract: Reinforcement learning has become a promising paradigm for improving large language model (LLM) agents on long-horizon search tasks, where the agent must make a sequence of intermediate decisions before receiving a final outcome. Ho…

  104. Hugging Face Daily Papers TIER_1 English(EN) ·

    面向Agentic强化学习的单回滚异步优化

    Asynchronous reinforcement learning with single-rollout optimization addresses stability issues in LLM training for complex tasks, outperforming existing methods in coding and reasoning benchmarks.

  105. Hugging Face Daily Papers TIER_1 English(EN) ·

    深度强化学习评估与设计范式的原则性分析

    Starting from the utilization of deep neural networks to approximate the state-action value function that led to winning one of the most challenging games, to algorithmic advancements that allowed solving problems without even explicitly stating the rules of the challenge at hand…

  106. arXiv cs.AI TIER_1 English(EN) · Xinbing Wang ·

    基于信息增益的推出策略优化:一种用于多轮 LLM Agent 的自适应树状推出方法

    Reinforcement learning has become a promising paradigm for improving large language model (LLM) agents on long-horizon search tasks, where the agent must make a sequence of intermediate decisions before receiving a final outcome. However, existing methods still face a key limitat…

  107. arXiv cs.AI TIER_1 English(EN) · Nijat Mehdiyev ·

    通过强化学习提升LLM生成过程模型质量:奖励函数设计的作用

    Large language models (LLMs) can generate BPMN process models from natural-language descriptions, yet supervised fine-tuning (SFT) limits their output quality to the patterns present in the training data. Reinforcement learning (RL) can optimize beyond this ceiling using external…

  108. arXiv cs.LG TIER_1 English(EN) · David Simchi-Levi ·

    多买方市场中的战略议价:基于可验证奖励的强化学习在LLM谈判中的应用

    Negotiation is a fundamental strategic interaction in management science, characterized by agents attempting to reach agreements while protecting private information, such as reservation costs and hidden valuations. A prevalent yet complex scenario involves a single seller negoti…

  109. arXiv cs.AI TIER_1 English(EN) · Faid Keddouri, Sohaib Houhou, Aissa Boulmerka, Nadir Farhi ·

    LLM增强的合作多智能体强化学习的政权条件稳定

    arXiv:2607.04470v1 Announce Type: cross Abstract: Large Language Models (LLMs) offer a natural interface for translating human objectives into reward signals for cooperative multi-agent reinforcement learning (MARL), yet the training-time dynamics of this integration remain poorl…

  110. arXiv cs.AI TIER_1 English(EN) · Weiyang Guo, Zesheng Shi, Longhui Zhang, Zeen Zhu, Min Zhang, Jing Li ·

    Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry

    arXiv:2607.03702v1 Announce Type: new Abstract: Large language model (LLM) agents have shown strong decision-making capabilities in long-horizon interactive tasks, yet they still struggle to effectively leverage failed trajectories: full retries incur high interaction costs, whil…

  111. arXiv cs.AI TIER_1 English(EN) · Mingxuan Fan, Peiyang Liu ·

    面向Agentic强化学习的面向进度和可靠性的群体策略优化

    arXiv:2607.04242v1 Announce Type: new Abstract: Group-based reinforcement learning (RL) has become an effective paradigm for improving large language model agents on long-horizon interactive tasks. To obtain finer-grained policy updates than trajectory-level optimization, recent …

  112. arXiv cs.AI TIER_1 English(EN) · Qiuyi Qi, Jinjian Zhang, Mutian Bao, Tian Liang, Guocong Li, Dongnan Liu, Wei Zhou, Jie Liu, Ming Kong, Linjian Mo, Feng Zhang, Qiang Zhu ·

    CARL: 具有约束意识的强化学习用于大型语言模型的规划

    arXiv:2607.04854v1 Announce Type: new Abstract: Despite their strong reasoning capabilities and extensive world knowledge, Large Language Models (LLMs) frequently generate plans that violate task constraints, undermining their reliability in real-world applications. This deficien…

  113. arXiv cs.AI TIER_1 English(EN) · Kai Zhao ·

    基于掩码的预测表示用于强化学习

    arXiv:2607.04153v1 Announce Type: cross Abstract: Vision-based deep reinforcement learning involves dealing with high-dimensional inputs of image information. It is crucial to abstract effective states from high-dimensional image inputs and limited samples for sample-efficient re…

  114. arXiv cs.AI TIER_1 English(EN) · Angen Ye, Weijie Ke, Xiaofeng Wang, Xinze Chen, Chaojun Ni, Guosheng Zhao, Boyuan Wang, Zheng Zhu, Junjie Xie, Dapeng Zhang ·

    HALO-WA: 用于世界-动作模型的混合注意力潜在引导在线强化学习

    arXiv:2607.04265v1 Announce Type: cross Abstract: World-action (WA) models can generate long-horizon action chunks for general-purpose robotic manipulation, but they remain vulnerable to calibration, perception, and contact-dynamics errors in real-world precision tasks, often fai…

  115. arXiv cs.AI TIER_1 English(EN) · Qiang Liu, Taian Guo, Ruizhi Qiao, Xing Sun ·

    RSPO:多轮LLM智能体的奖励交换策略优化

    arXiv:2607.04713v1 Announce Type: cross Abstract: Reinforcement learning holds significant potential for training large language models (LLMs) to handle multi-turn interactive tasks. However, in long-horizon, multi-turn tasks characterized by sparse outcome rewards, directly trai…

  116. arXiv cs.AI TIER_1 English(EN) · Saksham Sahai Srivastava, Vaneet Aggarwal ·

    面向大型语言模型强化学习技术的技术调查

    arXiv:2507.04136v2 Announce Type: replace Abstract: This survey offers a comprehensive foundation on the integration of RL with language models, highlighting prominent algorithms such as Proximal Policy Optimization (PPO), Q-Learning, and Actor-Critic methods. Additionally, it pr…

  117. arXiv cs.AI TIER_1 English(EN) · Siqi Zhu, Jiaxuan You ·

    OpenTinker: 在代理强化学习中分离关注点

    arXiv:2601.07376v2 Announce Type: replace Abstract: We introduce \textsc{OpenTinker}, an open infrastructure for training large language model (LLM) agents with many LoRA-backed policies over shared execution resources. Modern agent workloads mix supervised fine-tuning (SFT), onl…

  118. arXiv cs.AI TIER_1 English(EN) · Xiaoxuan Wang, Han Zhang, Haixin Wang, Yidan Shi, Ruoyan Li, Kaiqiao Han, Chenyi Tong, Haoran Deng, Renliang Sun, Alexander Taylor, Yanqiao Zhu, Jason Cong, Yizhou Sun, Wei Wang ·

    ARLArena:稳定智能体强化学习的统一框架

    arXiv:2602.21534v3 Announce Type: replace Abstract: Agentic reinforcement learning (ARL) has rapidly gained attention as a promising paradigm for training agents to solve complex, multi-step interactive tasks. Despite encouraging early results, ARL remains highly unstable, often …

  119. arXiv cs.AI TIER_1 English(EN) · David Schiff, Ofir Lindenbaum, Yonathan Efroni ·

    ICR-RL:通过上下文回归实现的深度强化学习

    arXiv:2509.11259v2 Announce Type: replace-cross Abstract: Recent advancements in machine learning have largely been driven by foundation models (FMs) trained on large, diverse datasets, enabling them to generalize effectively to new, related tasks. However, extending this paradig…

  120. arXiv cs.AI TIER_1 English(EN) · Lu Li, Tianwei Ni, Yihao Sun, Pierre-Luc Bacon ·

    离线到在线强化学习的三种模式

    arXiv:2510.01460v4 Announce Type: replace-cross Abstract: Offline-to-online reinforcement learning (RL) has emerged as a practical paradigm that leverages offline datasets for pretraining and online interactions for fine-tuning. However, its empirical behavior is highly inconsist…

  121. arXiv cs.AI TIER_1 English(EN) · Johannes Ackermann, Michael Noukhovitch, Takashi Ishida, Masashi Sugiyama ·

    梯度正则化缓解了来自人类反馈和可验证奖励的强化学习中的奖励破解问题

    arXiv:2602.18037v2 Announce Type: replace-cross Abstract: Reinforcement Learning from Human Feedback (RLHF) or Verifiable Rewards (RLVR) are two key steps in the post-training of modern Language Models (LMs). A common problem is reward hacking, where the policy may exploit inaccu…

  122. arXiv cs.AI TIER_1 English(EN) · Siyuan Wang, Lei Lei, Pranav Maheshwari, Sam Bellefeuille, Kan Zheng ·

    用于V2X资源分配的多智能体强化学习:通过基准测试解耦MARL挑战

    arXiv:2603.06607v2 Announce Type: replace-cross Abstract: Radio resource allocation (RRA) is a critical function in cellular vehicle-to-everything (C-V2X) networks, where vehicles must share limited wireless resources to support safety-critical communications. Multi-agent reinfor…

  123. arXiv cs.CL TIER_1 English(EN) · Mickel Liu, Liwei Jiang, Yancheng Liang, Simon Shaolei Du, Yejin Choi, Tim Althoff, Natasha Jaques ·

    使用在线自博弈强化学习追逐移动目标以实现更安全的语言模型

    arXiv:2506.07468v4 Announce Type: replace-cross Abstract: Conventional large language model (LLM) safety alignment relies on a reactive, disjoint loop: attackers exploit a static model, then defenders patch exposed vulnerabilities. This sequential setup leads to attackers overfit…

  124. arXiv cs.LG TIER_1 English(EN) · Georg Sch\"afer, Jakob Rehrl, Stefan Huber, Simon Hirlaender ·

    使用贝叶斯优化对能源感知强化学习进行样本高效帕累托前沿建模

    arXiv:2607.03140v1 Announce Type: new Abstract: Industrial automation increasingly demands control strategies that balance operational performance with strict energy efficiency requirements. A common approach to solving this multi-objective problem, particularly within the framew…

  125. arXiv cs.LG TIER_1 English(EN) · Jiayi Guan, Tianle Zhang, Li Shen, Ruiqi Zhang, Ao Zhou, Lusong Li, Guai Chen, Mengjie Li, Alois Knoll, Xiaodong He, Changjun Jiang ·

    CDCP:用于多任务离线安全强化学习的带上下文提示的条件扩散模型

    arXiv:2607.03903v1 Announce Type: new Abstract: Multi-task offline safe reinforcement learning (RL) promises to learn a shared optimal safe policy from offline data across multiple tasks. This paradigm provides an effective means for the widespread application of RL in multi-task…

  126. arXiv cs.LG TIER_1 English(EN) · Chee Heng Tan, Zhuoyi Lin, Mehul Motani, Wee Sun Lee ·

    关于奖励函数在大型语言模型置信度校准强化学习中的有效性

    arXiv:2607.04332v1 Announce Type: new Abstract: In this paper, we consider the setting where large language models (LLMs) are trained using reinforcement learning (RL) to simultaneously improve reasoning accuracy and verbalize its confidence. Our reward scheme uses two functions …

  127. arXiv cs.LG TIER_1 English(EN) · Mao-Lin Luo, Zhe-Xu Wang, Zi-Hao Zhou, Bo Ye, Jian Zhao, Min-Ling Zhang, Tong Wei ·

    RL 遗忘!迈向持续策略优化

    arXiv:2607.04364v1 Announce Type: new Abstract: Continual post-training is becoming a central paradigm for adapting vision-language models to evolving tasks. Recent work has increasingly favored reinforcement learning over supervised fine-tuning, driven by the belief that reinfor…

  128. arXiv cs.LG TIER_1 English(EN) · Fan Feng, Yujia Zheng, Minghao Fu, Yongqiang Chen, Guangyi Chen, Kevin Murphy, Biwei Huang, Kun Zhang ·

    通过协同式智能体探索和结构化建模学习任务充分的世界模型

    arXiv:2607.04409v1 Announce Type: new Abstract: Learning and planning in imagination using world models provides an effective paradigm for training agents for decision-making. However, existing approaches often rely on high-dimensional latent spaces or generic visual embeddings t…

  129. arXiv cs.LG TIER_1 English(EN) · Kyohei Suzuki, onstantinos Slavakis ·

    通过非单调包含实现非凸稀疏强化学习

    arXiv:2607.04990v1 Announce Type: new Abstract: This work delivers two key contributions: one to efficient feature selection in reinforcement learning (RL), the other to the theory of non-monotone inclusions. On the RL side, the estimation bias inherent in conventional regulariza…

  130. arXiv cs.LG TIER_1 English(EN) · Yujiang Li, Zhenyu Hou, Yi Jing, Jie Tang, Yuxiao Dong ·

    CompactionRL:具有上下文压缩的强化学习用于长视界智能体

    arXiv:2607.05378v1 Announce Type: new Abstract: Long-horizon agentic LLMs are increasingly limited by finite context windows, as extended interaction trajectories can exceed the maximum context length before a task is completed. Context compaction offers a natural solution by sum…

  131. arXiv cs.LG TIER_1 English(EN) · Jialun Cao, Fernando Acero, David \v{S}i\v{s}ka, Yufei Zhang ·

    熵正则化提高了连续时间强化学习中的策略鲁棒性

    arXiv:2607.03168v1 Announce Type: cross Abstract: Entropy regularization is widely used in continuous-time reinforcement learning (RL) to reduce sensitivity to environmental perturbations, yet its robustness benefits lack a rigorous theoretical foundation. This paper establishes …

  132. arXiv cs.LG TIER_1 English(EN) · Manuel Wendl, Lukas Koller, Tobias Ladner, Matthias Althoff ·

    使用基于集合的强化学习训练可验证鲁棒性智能体

    arXiv:2408.09112v2 Announce Type: replace Abstract: Reinforcement learning policies parametrized by deep neural networks have achieved strong performance for continuous control, yet even small input perturbations may lead to unpredictable behavior. This sensitivity limits their u…

  133. arXiv cs.LG TIER_1 English(EN) · Waris Radji, Thomas Michel, Hector Piteau ·

    Octax:JAX 中用于强化学习的加速 CHIP-8 街机环境

    arXiv:2510.01764v3 Announce Type: replace Abstract: Reinforcement learning (RL) research requires diverse, challenging environments that are both tractable and scalable. While modern video games may offer rich dynamics, they are computationally expensive and poorly suited for lar…

  134. arXiv cs.LG TIER_1 English(EN) · Vidur Sinha, Muhammed Ustaomeroglu, Guannan Qu ·

    基于Transformer的多智能体强化学习在具有长程交互的网络化系统中应用

    arXiv:2511.13103v2 Announce Type: replace Abstract: Multi-agent reinforcement learning (MARL) has shown promise for large-scale network control, yet existing methods face two major limitations. First, they typically rely on an exponential decay property of agent interactions on f…

  135. arXiv cs.LG TIER_1 English(EN) · Tianshi Xu, Yuteng Chen, Meng Li ·

    CLEANER:自净化轨迹增强了智能体强化学习

    arXiv:2601.15141v2 Announce Type: replace Abstract: Agentic Reinforcement Learning (RL) has empowered Large Language Models (LLMs) to utilize tools like Python interpreters for complex problem-solving. However, for parameter-constrained models (e.g., 4B--7B), the exploration phas…

  136. arXiv cs.LG TIER_1 English(EN) · Bang Giang Le, Viet Cuong Ta ·

    多智能体强化学习中的显式信用分配:基于局部奖励和依赖图

    arXiv:2601.21523v2 Announce Type: replace Abstract: To promote cooperation in Multi-Agent Reinforcement Learning, the reward signals of all agents can be aggregated together, forming global rewards that are commonly known as the fully cooperative setting. However, global rewards …

  137. arXiv cs.CL TIER_1 English(EN) · Jingjing Chen ·

    TurnOPD:使on-policy蒸馏具备回合感知能力,以实现高效长时域智能体训练

    On-policy distillation (OPD) trains a student policy by matching a stronger teacher on the student's own trajectories, offering a promising framework for language agent training. However, its application to long-horizon agentic tasks remains insufficiently explored. We identify t…

  138. Hugging Face Daily Papers TIER_1 English(EN) ·

    TurnOPD:使on-policy蒸馏具备回合感知能力,以实现高效长时域智能体训练

    Turn-level budgeting strategy for efficient on-policy distillation in long-horizon agent training addresses inefficiencies in full-horizon rollouts and shallow token concentration.

  139. Hugging Face Daily Papers TIER_1 English(EN) ·

    RoboTALES:通过任务对齐的模拟未来学习推理引导的机器人策略

    RoboTALES introduces a two-stage framework that combines LLM-based planning and VLM-based criticism to improve task-aligned video generation and robotic policy training.

  140. arXiv cs.LG TIER_1 English(EN) · Yuxiao Dong ·

    CompactionRL:具有上下文压缩的强化学习用于长视野智能体

    Long-horizon agentic LLMs are increasingly limited by finite context windows, as extended interaction trajectories can exceed the maximum context length before a task is completed. Context compaction offers a natural solution by summarizing previous interaction states and continu…

  141. arXiv cs.LG TIER_1 English(EN) · onstantinos Slavakis ·

    非凸稀疏强化学习通过非单调包含

    This work delivers two key contributions: one to efficient feature selection in reinforcement learning (RL), the other to the theory of non-monotone inclusions. On the RL side, the estimation bias inherent in conventional regularization schemes is addressed by augmenting classica…

  142. Hugging Face Daily Papers TIER_1 English(EN) ·

    CARL: 具有约束意识的强化学习用于大型语言模型的规划

    Despite their strong reasoning capabilities and extensive world knowledge, Large Language Models (LLMs) frequently generate plans that violate task constraints, undermining their reliability in real-world applications. This deficiency arises from a lack of systematic mechanisms t…

  143. arXiv cs.AI TIER_1 English(EN) · Qiang Zhu ·

    CARL:LLM规划的约束感知强化学习

    Despite their strong reasoning capabilities and extensive world knowledge, Large Language Models (LLMs) frequently generate plans that violate task constraints, undermining their reliability in real-world applications. This deficiency arises from a lack of systematic mechanisms t…

  144. Hugging Face Daily Papers TIER_1 English(EN) ·

    使用离线强化学习控制LLM Agent的工具链

    Large language model (LLM) agents are usually improved by changing prompts, models, or hand-written workflows, while the execution harness around the model is treated as fixed infrastructure. We argue that this harness is itself a learnable control layer. We formalize harness ope…

  145. arXiv cs.LG TIER_1 English(EN) · Kevin Wang, Kevin Yang, Arjun Prakash, Amy Greenwald ·

    面向双人零和不完美信息博弈策略表示的学习

    arXiv:2607.01498v1 Announce Type: new Abstract: We investigate the problem of learning useful policy representations (embeddings) in two-player zero-sum imperfect-information games. We make three contributions: First, we introduce methods of creating datasets of policies for a gi…

  146. arXiv cs.LG TIER_1 English(EN) · Jen-Yen Chang, Takayuki Osa, Tatsuya Harada ·

    强化学习中分类批评者支持的学习

    arXiv:2607.01880v1 Announce Type: new Abstract: Value functions are an essential component in actor-critic based deep reinforcement learning (RL). Conventionally, these functions are trained as a regression task by minimising the mean squared error (MSE) relative to bootstrapped …

  147. arXiv cs.AI TIER_1 English(EN) · Yilie Huang, Wenpin Tang, Xun Yu Zhou ·

    ART for Diffusion Sampling: Continuous-Time Control and Actor-Critic Learning

    arXiv:2607.02137v1 Announce Type: cross Abstract: We study timestep allocation for score-based diffusion sampling, where a learned reverse-time dynamics is discretized on a finite grid. Uniform and hand-crafted schedules are standard choices, but they rely on fixed prescriptions …

  148. arXiv cs.AI TIER_1 English(EN) · Yuriy Maksyuta, George Bredis, Ruslan Rakhimov, Daniil Gavrilov ·

    Rank-Then-Act:无奖励的帧序进展控制

    arXiv:2607.01897v1 Announce Type: cross Abstract: We introduce Rank-Then-Act (RTA), a framework for learning control policies from expert video demonstrations without environment rewards. RTA trains a Vision-Language Model (VLM) offline as a progress-based ordinal scorer, using a…

  149. arXiv cs.AI TIER_1 English(EN) · Shenghui Zhang, YuXuan Gao, Songwei Zhao, Jifeng Hu, Zijing Zhang, Hechang Chen ·

    面向端到端无人机导航的轻量级安全强化学习

    arXiv:2607.01794v1 Announce Type: cross Abstract: With the rapid development of autonomous aerial systems, Unmanned Aerial Vehicles (UAVs) are increasingly deployed in applications such as inspection, environmental monitoring, and rescue, creating growing demand for reliable auto…

  150. arXiv cs.AI TIER_1 English(EN) · Juliette Decugis, Sean O'Brien, Francis Bach, Gabriel Synnaeve, Taco Cohen ·

    别让收益消退:解析强化学习中的策略梯度权重

    arXiv:2607.01490v1 Announce Type: cross Abstract: Reinforcement learning post-training dramatically improves LLM reasoning, but suffers from training instability and diversity collapse. Advantage functions offer an appealing fix: they reshape the training objective, reweight whic…

  151. arXiv cs.LG TIER_1 English(EN) · Ren\'e Carmona, Mathieu Lauri\`ere ·

    均场强化学习

    arXiv:2607.01525v1 Announce Type: cross Abstract: This monograph provides an introduction to mean field reinforcement learning through the lens of Markov decision processes arising from large-population stochastic control with mean field interactions and common noise. Starting fr…

  152. arXiv cs.AI TIER_1 English(EN) · Anish Diwan, Davide Tateo, Christopher E. Mower, Haitham Bou-Ammar, Jan Peters, Oleg Arenz ·

    Trust Region Inverse Reinforcement Learning: Explicit Dual Ascent using Local Policy Updates

    arXiv:2605.11020v2 Announce Type: replace-cross Abstract: Inverse reinforcement learning (IRL) is typically formulated as maximizing entropy subject to matching the distribution of expert trajectories. Classical (dual-ascent) IRL guarantees monotonic performance improvement but r…

  153. arXiv cs.LG TIER_1 English(EN) · Daniel Thi Graviet, Lovre Pesut, Ivan Dagelic, Vedran Jukic, Ivan Burazin ·

    编码智能体强化学习中的部署基础设施税

    arXiv:2607.01415v1 Announce Type: new Abstract: Coding-agent reinforcement learning treats execution infrastructure as a background implementation detail, despite relying on large numbers of interactive software rollouts. This is a missed opportunity: measuring infrastructure ove…

  154. arXiv cs.AI TIER_1 English(EN) · Xun Yu Zhou ·

    ART for Diffusion Sampling: Continuous-Time Control and Actor-Critic Learning

    We study timestep allocation for score-based diffusion sampling, where a learned reverse-time dynamics is discretized on a finite grid. Uniform and hand-crafted schedules are standard choices, but they rely on fixed prescriptions and can therefore be suboptimal. To address this l…

  155. Hugging Face Daily Papers TIER_1 English(EN) ·

    ART for Diffusion Sampling: Continuous-Time Control and Actor-Critic Learning

    We study timestep allocation for score-based diffusion sampling, where a learned reverse-time dynamics is discretized on a finite grid. Uniform and hand-crafted schedules are standard choices, but they rely on fixed prescriptions and can therefore be suboptimal. To address this l…

  156. arXiv cs.LG TIER_1 English(EN) · Jinwen Wang, Youfang Lin, Xiaobo Hu, Shuo Wang, Kai Lv ·

    局部运动至关重要:用于视频强化学习预训练的解构-重构范式

    arXiv:2607.00808v1 Announce Type: new Abstract: Pre-training on large-scale videos to improve reinforcement learning efficiency is promising yet remains challenging. Existing methods typically treat the agent as an indivisible entity, modeling motion patterns globally. Such globa…

  157. arXiv cs.AI TIER_1 English(EN) · Zishang Jiang, Jinyi Han, Tingyun Li, Xinyi Wang, Sihang Jiang, Jiaqing Liang, Zhaoqian Dai, Shuguang Ma, Fei Yu, Yanghua Xiao ·

    LLM强化学习中有效且多样化探索的选择性专家指导

    arXiv:2510.04140v2 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become a widely adopted technique for enhancing the reasoning ability of Large Language Models (LLMs). However, the effectiveness of RLVR strongly depends on the capabili…

  158. arXiv cs.AI TIER_1 English(EN) · Zhiqi Li, Wen Zhang, Bo Zhu ·

    Flow-Map GRPO:基于锚定随机组合的少样本流图生成器的强化学习

    arXiv:2607.00535v1 Announce Type: cross Abstract: Few-step flow-map generators, such as consistency models and MeanFlow, accelerate sampling by directly learning long-range transport maps between noise and data. However, these models are typically deterministic, which makes them …

  159. arXiv cs.AI TIER_1 English(EN) · Konstantin Garbers ·

    在Actor-Critic强化学习中评估、测量和控制Critic的复杂性

    arXiv:2607.00452v1 Announce Type: cross Abstract: Actor-critic methods depend on learned critics, but critic quality is often evaluated only indirectly through return, temporal-difference error, or value loss. Critic complexity is introduced as an additional diagnostic and interv…

  160. arXiv cs.AI TIER_1 English(EN) · Jongchan Park, Seungjun Oh, Seungho Baek, Yusung Kim ·

    通过数据高效的无监督强化学习学习可泛化技能策略

    arXiv:2607.00392v1 Announce Type: cross Abstract: Unsupervised Reinforcement Learning (URL) aims to pre-train scalable, skill-conditioned policies without extrinsic rewards, serving as a foundation for downstream control tasks. Despite recent progress, we argue that current off-p…

  161. arXiv cs.AI TIER_1 English(EN) · Egor Cherepanov, Daniil Zelezetsky, Alexey K. Kovalev, Aleksandr I. Panov ·

    KAGE-Bench:强化学习的快速已知轴视觉泛化评估

    arXiv:2601.14232v2 Announce Type: replace-cross Abstract: Pixel-based reinforcement learning agents often fail under purely visual distribution shift even when latent dynamics and rewards are unchanged, but existing benchmarks entangle multiple sources of shift and hinder systema…

  162. arXiv cs.LG TIER_1 English(EN) · Gaia Molinaro, Anne G. E. Collins ·

    奖励函数压缩促进目标相关强化学习

    arXiv:2509.06810v3 Announce Type: replace-cross Abstract: Humans can uniquely assign value to novel, abstract outcomes to support reinforcement learning. However, this flexibility is cognitively costly and reduces learning efficiency. We propose that goal-dependent learning initi…

  163. arXiv cs.CL TIER_1 English(EN) · Ziyou Hu, Zhengliang Shi, Minghang Zhu, Haitao Li, Teng Sun, Pengjie Ren, Suzan Verberne, Zhaochun Ren ·

    OpenReward:通过强化学习学习奖励长篇代理任务

    arXiv:2510.24636v3 Announce Type: replace Abstract: Reward models (RMs) have become essential for aligning large language models (LLMs), serving as scalable proxies for human evaluation in both training and inference. However, existing RMs struggle on knowledge-intensive and long…

  164. arXiv cs.LG TIER_1 English(EN) · Jinwen Wang, Youfang Lin, Xiaobo Hu, Qian Xu, Shuo Wang, Zhuo Chen, Kai Lv ·

    视觉强化学习泛化任务相关表示解耦

    arXiv:2607.00796v1 Announce Type: new Abstract: Visual Reinforcement Learning (VRL) has achieved considerable success in solving control tasks. However, generalizing learned policies to new environments remains a major challenge, as agents often overfit to task-irrelevant feature…

  165. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Mathieu Laurière ·

    均场强化学习

    This monograph provides an introduction to mean field reinforcement learning through the lens of Markov decision processes arising from large-population stochastic control with mean field interactions and common noise. Starting from the connection between multi-agent reinforcemen…

  166. arXiv cs.LG TIER_1 English(EN) · Kai Lv ·

    局部运动很重要:用于视频强化学习预训练的解构-重构范式

    Pre-training on large-scale videos to improve reinforcement learning efficiency is promising yet remains challenging. Existing methods typically treat the agent as an indivisible entity, modeling motion patterns globally. Such global modeling is tightly coupled with the morpholog…

  167. arXiv cs.LG TIER_1 English(EN) · Kai Lv ·

    视觉强化学习泛化任务相关表示解耦

    Visual Reinforcement Learning (VRL) has achieved considerable success in solving control tasks. However, generalizing learned policies to new environments remains a major challenge, as agents often overfit to task-irrelevant features in the training environment. To solve this pro…

  168. arXiv cs.AI TIER_1 English(EN) · Bo Zhu ·

    Flow-Map GRPO:基于锚定随机组合的少样本流图生成器的强化学习

    Few-step flow-map generators, such as consistency models and MeanFlow, accelerate sampling by directly learning long-range transport maps between noise and data. However, these models are typically deterministic, which makes them difficult to optimize with reinforcement learning …

  169. arXiv cs.AI TIER_1 English(EN) · Konstantin Garbers ·

    在Actor-Critic强化学习中评估、测量和控制Critic的复杂性

    Actor-critic methods depend on learned critics, but critic quality is often evaluated only indirectly through return, temporal-difference error, or value loss. Critic complexity is introduced as an additional diagnostic and intervention dimension for actor-critic reinforcement le…

  170. arXiv cs.AI TIER_1 English(EN) · Debmalya Mandal, Andi Nika, Parameswaran Kamalaruban, Adish Singla, Goran Radanovi\'c ·

    Corruption Robust Offline Reinforcement Learning with Human Feedback

    arXiv:2402.06734v2 Announce Type: replace-cross Abstract: We study data corruption robustness for reinforcement learning with human feedback (RLHF) in an offline setting. Given an offline dataset of pairs of trajectories along with feedback about human preferences, an $\varepsilo…

  171. arXiv cs.AI TIER_1 English(EN) · Yuanda Xu, Zhengze Zhou, Hejian Sang, Xiaomin Li, Jiaxin Zhang, Xinchen Du, Zhipeng Wang, Alborz Geramifard ·

    TRIAGE:用于代理强化学习的角色类型信用分配

    arXiv:2606.32017v1 Announce Type: cross Abstract: Agentic reinforcement learning requires assigning credit to environment-facing actions such as searches, clicks, edits, navigation commands, and object interactions. Standard GRPO uses the final verifier outcome as a uniform advan…

  172. arXiv cs.AI TIER_1 English(EN) · Arshia Rafieioskouei, Tzu-Han Hsu, Matthew Lucas, Borzoo Bonakdarpour ·

    HyPOLE:部分可观察性下超属性引导的多智能体强化学习

    arXiv:2606.30966v1 Announce Type: new Abstract: Formal specification is a powerful tool to guide the learning process and provides significant advantages over reward shaping: (1) mathematical rigor; (2) expressiveness to specify objectives and constraints, and (3) the ability to …

  173. arXiv cs.LG TIER_1 English(EN) · Deepak Akhare, Luning Sun, Xin-Yang Liu, Xiantao Fan, Timo Bremer, Ben Zhu, Jian-Xun Wang ·

    面向流体控制的离线强化学习:基于数据的多观测策略提取

    arXiv:2606.31025v1 Announce Type: new Abstract: Active flow control is a fundamental application in engineering. Recent advances in deep reinforcement learning have made progress in this field. However, the classical online RL approaches require extensive real-time interactions w…

  174. arXiv cs.LG TIER_1 English(EN) · Mingyi Li, Taira Tsuchiya, Kenji Yamanishi ·

    策略优化在未知转移的MDP中实现数据依赖的遗憾界限

    arXiv:2606.31769v1 Announce Type: new Abstract: We study policy optimization for online episodic tabular Markov decision processes with unknown transition kernels, aiming for best-of-both-worlds guarantees together with data-dependent regret bounds. Recent work (Dann et al., 2023…

  175. arXiv cs.AI TIER_1 English(EN) · Yang Yang, Bingjie Chen, Zihan Wang, Yizhe Li, Guoping Pan, Yi Cheng, Houde Liu ·

    Stage-Transition Dense Reward Modeling for Reinforcement Learning

    arXiv:2606.31377v1 Announce Type: cross Abstract: Reinforcement learning for long-horizon robotic manipulation is often limited by sparse and delayed rewards, while manually designing dense shaping signals is costly and brittle to changes in environments and object configurations…

  176. arXiv cs.LG TIER_1 English(EN) · Ethan Hirschowitz, Fabio Ramos ·

    Warp RL:重塑基础策略分布以适应动态变化

    arXiv:2606.31043v1 Announce Type: new Abstract: Residual reinforcement learning adapts a pretrained robot policy by learning an additive correction to its actions. While effective when adaptation amounts to shifting the base policy's action distribution, additive corrections cann…

  177. arXiv cs.LG TIER_1 English(EN) · Tommaso Marzi, Cesare Alippi, Andrea Cini ·

    多智能体强化学习的分层消息传递策略

    arXiv:2507.23604v2 Announce Type: replace Abstract: Decentralized Multi-Agent Reinforcement Learning (MARL) methods allow for learning scalable multi-agent policies, but suffer from partial observability and induced non-stationarity. These challenges can be addressed by introduci…

  178. arXiv cs.LG TIER_1 English(EN) · Yusung Kim ·

    通过数据高效的无监督强化学习学习可泛化技能策略

    Unsupervised Reinforcement Learning (URL) aims to pre-train scalable, skill-conditioned policies without extrinsic rewards, serving as a foundation for downstream control tasks. Despite recent progress, we argue that current off-policy URL methods are limited by two critical, ove…

  179. arXiv cs.AI TIER_1 English(EN) · Alborz Geramifard ·

    TRIAGE:用于代理强化学习的角色类型信用分配

    Agentic reinforcement learning requires assigning credit to environment-facing actions such as searches, clicks, edits, navigation commands, and object interactions. Standard GRPO uses the final verifier outcome as a uniform advantage over all action tokens. This outcome signal i…

  180. arXiv cs.LG TIER_1 English(EN) · Kenji Yamanishi ·

    策略优化在未知转移的MDP中实现数据依赖的遗憾界限

    We study policy optimization for online episodic tabular Markov decision processes with unknown transition kernels, aiming for best-of-both-worlds guarantees together with data-dependent regret bounds. Recent work (Dann et al., 2023; Li et al., 2026) has shown that policy optimiz…

  181. arXiv cs.AI TIER_1 English(EN) · Qijun Li, Zheng Fu, Qi Song, Yifei He, Weitao Zhou, Kun Jiang, Diange Yang ·

    具有状态感知的探索的双流强化学习

    arXiv:2606.29820v1 Announce Type: cross Abstract: In complex continuous-control reinforcement learning tasks, multimodal optimal actions often coincide with uncertain, multimodal return distributions, making reliable value estimation and multimodal exploration challenging. Existi…

  182. arXiv cs.AI TIER_1 English(EN) · Runze Zhao, Dongruo Zhou, Sumit Kumar Jha, Nathaniel D. Bastian, Ankit Shah ·

    HiComm:多智能体强化学习的层次化通信

    arXiv:2606.29126v1 Announce Type: new Abstract: Cooperative multi-agent reinforcement learning (MARL) often relies on communication to mitigate partial observability, yet most existing protocols treat messages as flat dense vectors detached from the structure of the observations …

  183. arXiv cs.AI TIER_1 English(EN) · Daiki E. Matsunaga, Junho Na, Tri Wahyu Guntara, Scott Sanner, Pascal Poupart, Jongmin Lee, Kee-Eung Kim ·

    ACPO:多智能体强化学习的智能体链式策略优化

    arXiv:2606.30072v1 Announce Type: new Abstract: Cooperative tasks in Multi-Agent Reinforcement Learning (MARL) require agents to collectively maximize a shared return. Under the Centralized Training with Decentralized Execution (CTDE) paradigm, policy gradients have remained diff…

  184. arXiv cs.AI TIER_1 English(EN) · Adithya Mohan, Daniel Kriegl, Torsten Sch\"on ·

    RoAd-RL:用于鲁棒对抗强化学习的统一库和基准测试

    arXiv:2606.29867v1 Announce Type: cross Abstract: Deep Reinforcement Learning (DRL) has achieved significant success in robotics and autonomous systems, yet remains vulnerable to adversarial perturbations that can severely degrade performance. Research in adversarial reinforcemen…

  185. arXiv cs.AI TIER_1 English(EN) · Zibin Meng, Kani Chen ·

    CRAFT: 来自自由同级回放的逆事实信用分配,用于自蒸馏代理强化学习

    arXiv:2606.29476v1 Announce Type: cross Abstract: Self-distilled agentic reinforcement learning augments trajectory-level reward with a token-level distillation loss, using as its teacher the same policy conditioned on privileged context. The prevailing recipe gates this loss by …

  186. arXiv cs.LG TIER_1 English(EN) · Zheming Fu, Ruizhe He, Wei Shang, Xiaoxiao Ma, Lei Wang, Chang Liu, Siming Fu ·

    FlowAWR:通过优势加权校正实现在线自适应流强化

    arXiv:2606.30376v1 Announce Type: new Abstract: Aligning generative flow models on continuous spaces via online reinforcement learning is constrained by intractable trajectory likelihoods. Existing density-approximated policy gradient methods rely on stochastic SDE samplers to co…

  187. arXiv cs.LG TIER_1 English(EN) · Jing Liang, Hongyao Tang, Yi Ma, Yancheng He, Weixun Wang, Xiaoyang Li, Ju Huang, Wenbo Su, Jinyi Liu, Yan Zheng, Jianye Hao, Bo Zheng ·

    优化训练策略的幻象:单调推理策略作为 LLM 强化学习的真正目标

    arXiv:2606.29526v1 Announce Type: new Abstract: Reinforcement learning (RL) has gained growing attention in large language model (LLM) post-training, yet RL training remains fragile and can suffer from instability or collapse. One vital cause is training-inference mismatch: LLM a…

  188. arXiv cs.LG TIER_1 English(EN) · Jesse Ponnock, Lucas Ho ·

    强化学习在《超级马力欧兄弟》中的应用:1-1关卡的课程设计、教学法与最优关卡设计

    arXiv:2606.29511v1 Announce Type: new Abstract: World 1-1 of Super Mario Bros is widely celebrated as a masterclass in game design: its progressive structure is credited with teaching players core mechanics through the level itself. We ask whether that structure is empirically me…

  189. arXiv cs.AI TIER_1 English(EN) · Hao Wang, Jiuzhou Lei, Dayou Li, Bangya Liu, Minghui Zheng, Manling Li, Ruohan Zhang, Zhiwen Fan ·

    行为去克隆:在无推理时转向的情况下将模式重定向蒸馏到策略权重中

    arXiv:2606.29201v1 Announce Type: cross Abstract: Behavior-cloned policies often learn multiple behavior modes from demonstration datasets, including modes that are unsafe or otherwise undesired at deployment. For example, a policy trained on diverse handover demonstrations may l…

  190. arXiv cs.LG TIER_1 English(EN) · Amritansh Mishra, Supriyo Chakraborty, Berkcan Kapusuzoglu ·

    关于群组相对策略优化策略梯度基础:信用分配、梯度稀疏性和秩崩溃

    arXiv:2606.29238v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) eliminates the learned critic in PPO by using the mean reward of grouped rollouts as a baseline. We provide a rigorous derivation of GRPO from first principles of the policy gradient theorem…

  191. arXiv cs.LG TIER_1 English(EN) · Congde Hu, Danping Li, Lin Xu, Wenying Xu ·

    面向状态切换扩散模型中线性二次Stackelberg微分博弈的熵正则化强化学习

    arXiv:2606.28671v1 Announce Type: new Abstract: Stackelberg differential games (SDGs) provide a powerful framework for hierarchical decision-making in stochastic and continuous-time environments, yet their solution remains computationally challenging due to the complexity of trad…

  192. arXiv cs.LG TIER_1 English(EN) · Congde Hu, Zhuo Jin, Danping Li, Lin Xu ·

    基于熵正则化强化学习的切换状态跳扩散过程零和随机微分博弈

    arXiv:2606.28669v1 Announce Type: new Abstract: To address parameter misspecification and sudden structural environmental changes in conventional stochastic differential game (SDG) frameworks, this paper introduces a distributional control approach that characterizes optimal stra…

  193. arXiv cs.AI TIER_1 English(EN) · Rajib Mostakim, Reza T. Batley, Sourav Saha ·

    通过可分离神经网络架构和应用实现敏捷强化学习

    arXiv:2601.23225v2 Announce Type: replace-cross Abstract: Deep reinforcement learning (RL) is increasingly deployed in resource-constrained environments, yet go-to function approximators - multilayer perceptrons (MLPs) - are often parameter-inefficient due to an imperfect inducti…

  194. arXiv cs.AI TIER_1 English(EN) · Mathieu Petitbois, R\'emy Portelas, Sylvain Lamprier ·

    鲁棒风格对齐下高质量行为的离线强化学习

    arXiv:2601.22823v2 Announce Type: replace-cross Abstract: We study offline reinforcement learning of style-conditioned policies using explicit style supervision via subtrajectory labeling functions. In this setting, aligning style with high task performance is particularly challe…

  195. arXiv cs.AI TIER_1 English(EN) · Debmalya Mandal, Paulius Sasnauskas, Goran Radanovic ·

    基于人类反馈的分布鲁棒强化学习

    arXiv:2503.00539v2 Announce Type: replace-cross Abstract: Reinforcement learning from human feedback (RLHF) has evolved to be one of the main methods for fine-tuning large language models (LLMs). However, existing RLHF methods are non-robust, and their performance deteriorates if…

  196. arXiv cs.AI TIER_1 English(EN) · Charles Westphal, Stephen Hailes, Mirco Musolesi ·

    TERC:强化学习中状态变量选择的迁移熵冗余准则

    arXiv:2401.11512v2 Announce Type: replace-cross Abstract: Identifying the most suitable variables to represent the state is a fundamental challenge in Reinforcement Learning (RL). These variables must efficiently capture the information necessary for making optimal decisions. In …

  197. Hugging Face Daily Papers TIER_1 English(EN) ·

    TRIAGE:用于代理强化学习的角色类型信用分配

    TRIAGE introduces a role-typed credit assignment framework that enhances agentic reinforcement learning by providing more nuanced credit assignment than standard GRPO methods.

  198. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Borzoo Bonakdarpour ·

    HyPOLE:部分可观测性下超属性引导的多智能体强化学习

    Formal specification is a powerful tool to guide the learning process and provides significant advantages over reward shaping: (1) mathematical rigor; (2) expressiveness to specify objectives and constraints, and (3) the ability to define tactics to achieve objectives. However, t…

  199. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Claudio Pacchierotti ·

    基于采样的协调感知多目标多机器人强化学习

    Multi-robot systems must simultaneously optimize competing objectives while maintaining coordinated behavior. Existing multi-agent reinforcement learning approaches often rely on fixed or centralized coordination, which limits adaptability and violates distributed constraints. Th…

  200. arXiv cs.LG TIER_1 English(EN) · Siming Fu ·

    FlowAWR:通过优势加权校正实现在线自适应流强化

    Aligning generative flow models on continuous spaces via online reinforcement learning is constrained by intractable trajectory likelihoods. Existing density-approximated policy gradient methods rely on stochastic SDE samplers to construct tractable transition kernels, which intr…

  201. arXiv cs.AI TIER_1 English(EN) · Difan Jiao, Raghav Singhal, Robert West, Ashton Anderson ·

    Tandem Reinforcement Learning with Verifiable Rewards

    arXiv:2606.28166v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has significantly improved the reasoning capability of large language models, reaching expert or even superhuman performance in domains such as competition math. However, whether…

  202. arXiv cs.AI TIER_1 English(EN) · Qitai Tan, Zefang Zong, Yang Li, Peng Chen ·

    ATOD:用于多轮自主代理的退火转弯感知策略内蒸馏

    arXiv:2606.27814v1 Announce Type: new Abstract: Training small language-model agents for long-horizon interactive tasks requires both fast imitation and reward-driven improvement. On-policy distillation (OPD) provides dense teacher guidance and typically improves rapidly in the e…

  203. arXiv cs.LG TIER_1 English(EN) · Jiexin Wang, Eiji Uchibe ·

    正则化奖惩强化学习

    arXiv:2606.28152v1 Announce Type: new Abstract: We propose KL-Coupled Policy Regularization (KCPR), a policy coordination framework for Reward-Punishment Reinforcement Learning (RPRL). Based on KCPR, we derive KL-Coupled Soft Optimality (KCSO) and develop its deep realization, kl…

  204. arXiv cs.LG TIER_1 English(EN) · Tianlin Pan, Lianyu Pang, Cheng Da, Huan Yang, Changqian Yu, Kun Gai, Wenhan Luo ·

    NormGuard:流匹配强化学习中的保奖励范数约束

    arXiv:2606.27771v1 Announce Type: new Abstract: Reinforcement learning (RL) post-training improves the reward alignment of flow-based generators, but often degrades perceptual quality in ways that are not captured by the reward proxy. We identify a simple structural signature of …

  205. arXiv cs.AI TIER_1 English(EN) · Austin A. Nguyen, Michael P. Wellman ·

    离线博弈论多智能体强化学习中的保守均衡发现

    arXiv:2603.00374v2 Announce Type: replace Abstract: Offline learning of strategies takes data efficiency to its extreme by restricting algorithms to a fixed dataset of state-action trajectories. We consider the problem in a mixed-motive multiagent setting, where the goal is to so…

  206. Hugging Face Daily Papers TIER_1 English(EN) ·

    优化训练策略的幻觉:单调推理策略作为 LLM 强化学习的真正目标

    Training-inference mismatch in reinforcement learning for large language models leads to instability, which is addressed through a new policy optimization objective and framework that ensures consistent policy improvements between training and inference phases.

  207. arXiv cs.AI TIER_1 English(EN) · Ashton Anderson ·

    Tandem Reinforcement Learning with Verifiable Rewards

    Reinforcement learning with verifiable rewards (RLVR) has significantly improved the reasoning capability of large language models, reaching expert or even superhuman performance in domains such as competition math. However, whether weaker agents and humans can actually harness t…

  208. arXiv cs.LG TIER_1 English(EN) · Eiji Uchibe ·

    正则化奖惩强化学习

    We propose KL-Coupled Policy Regularization (KCPR), a policy coordination framework for Reward-Punishment Reinforcement Learning (RPRL). Based on KCPR, we derive KL-Coupled Soft Optimality (KCSO) and develop its deep realization, klDMP. Unlike existing RPRL approaches that optimi…

  209. arXiv cs.AI TIER_1 English(EN) · Peng Chen ·

    ATOD:用于多轮自主代理的退火转弯感知策略内蒸馏

    Training small language-model agents for long-horizon interactive tasks requires both fast imitation and reward-driven improvement. On-policy distillation (OPD) provides dense teacher guidance and typically improves rapidly in the early stage, but its gains saturate once the stud…

  210. arXiv cs.LG TIER_1 English(EN) · Wenhan Luo ·

    NormGuard:流匹配强化学习中的保奖励范数约束

    Reinforcement learning (RL) post-training improves the reward alignment of flow-based generators, but often degrades perceptual quality in ways that are not captured by the reward proxy. We identify a simple structural signature of this drift: across three post-training methods (…

  211. arXiv cs.LG TIER_1 English(EN) · Wenjie Huang, Yang Li, Jingjia Teng, Mingwei Jin, Kai Song, Yougang Bian, Yongfu Li, Qisong Yang, Helai Huang ·

    通过自适应奖励塑形和策略比重加权策略实现样本高效的迁移强化学习

    arXiv:2606.26527v1 Announce Type: new Abstract: Transfer learning improves policy learning efficiency by reusing knowledge from source tasks, providing a feasible paradigm for safe and efficient autonomous highway lane changing decision-making. Existing methods frequently encount…

  212. arXiv cs.LG TIER_1 English(EN) · Behnam Gheshlaghi, Bahador Rashidi, Shahin Atakishiyev ·

    Mesh-RL:耦合子网格强化学习

    arXiv:2606.26333v1 Announce Type: new Abstract: Reinforcement learning in large or sparse-reward environments suffers from slow temporal-difference reward propagation, as value information spreads only locally across the state space. We propose Mesh-RL, a spatial domain-decomposi…

  213. arXiv cs.AI TIER_1 English(EN) · Aniruddha Joshi, Niklas Lauffer, Sanjit Seshia ·

    多目标强化学习的确定性帕累托最优策略合成

    arXiv:2606.26397v1 Announce Type: cross Abstract: Real-world decision-making often requires balancing multiple conflicting objectives, a challenge that standard Reinforcement Learning (RL) frequently addresses by aggregating rewards into a single scalar signal. While effective fo…

  214. arXiv cs.AI TIER_1 English(EN) · Ting Zhou, Zhenqing Ling, Yiyang Zhao, Ying Shen, Daoyuan Chen ·

    GEOALIGN:用于鲁棒LLM强化学习的几何回滚策展

    arXiv:2606.26917v1 Announce Type: cross Abstract: Online reinforcement learning is widely used to align large language models (LLMs) with reward signals, yet training can be unstable under noisy or misspecified rewards. We identify a failure mode we call directional inconsistency…

  215. arXiv cs.AI TIER_1 English(EN) · Jesper Klicks, Sander Vr\v{z}ina, Vincent Fran\c{c}ois-Lavet ·

    状态表示在深度强化学习中的重要性:以能源交易为例

    arXiv:2606.27032v1 Announce Type: cross Abstract: Energy trading decisions depend not only on current market prices, but also on expected future market conditions, and operational constraints. This makes the state representation given to a reinforcement learning agent an importan…

  216. arXiv cs.CL TIER_1 English(EN) · Shuo Yang, Jinyang Wu, Zhengxi Lu, Yuhao Shen, Fan Zhang, Lang Feng, Shuai Zhang, Haoran Luo, Zheng Lian, Zhengqi Wen, Jianhua Tao ·

    OPID:用于智能体强化学习的策略内技能蒸馏

    arXiv:2606.26790v1 Announce Type: new Abstract: Outcome-based reinforcement learning provides a stable optimization backbone for language agents, but its sparse trajectory-level rewards provide little guidance on which intermediate decisions should be reinforced or suppressed. On…

  217. arXiv cs.LG TIER_1 English(EN) · Rongjian Chen, Jianmin Hu, Kejiang Ye, Minxian Xu ·

    RolloutPipe:在分离式策略内强化学习大模型中实现重叠流水线式部署和训练

    arXiv:2606.26997v1 Announce Type: cross Abstract: Large language model (LLM) post-training for reasoning increasingly relies on reinforcement learning with verifiable rewards (RLVR), where models learn from ground-truth feedback on mathematical, logical, and scientific tasks. To …

  218. arXiv cs.LG TIER_1 English(EN) · Erhan Bayraktar, Martin Hernandez, Qinxin Yan, Yuhua Zhu ·

    Mean-Field PhiBE:从离散时间数据中进行连续时间均值场强化学习

    arXiv:2606.26498v1 Announce Type: cross Abstract: This paper addresses model-free continuous-time mean-field control in a setting where the population dynamics evolve continuously according to an unknown McKean-Vlasov stochastic differential equation, while only discrete-time tra…

  219. arXiv cs.LG TIER_1 English(EN) · Yingyu Lin, Qiyue Gao, Nikki Lijing Kuang, Xunpeng Huang, Kun Zhou, Tongtong Liang, Zhewei Yao, Yi-An Ma, Yuxiong He ·

    无需真实解的强化学习可改进LLM

    arXiv:2606.27369v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) for training LLMs typically rely on ground-truth answers to assign rewards, limiting their applicability to tasks where the ground-truth solution is unknown. We introduce a \text…

  220. arXiv cs.AI TIER_1 English(EN) · Shicheng Ye, Chao Yu ·

    大型语言模型智能体的经验规则与策略联合学习

    arXiv:2606.27136v1 Announce Type: new Abstract: For LLM agents in multi-step interactive environments, a key challenge is to make effective use of accumulated interaction experience. Existing work has typically separated two uses of such experience: keeping it outside the model a…

  221. arXiv cs.AI TIER_1 English(EN) · Boyun Zhang, Chao Wang, Kai Wu ·

    EVOM: 强化学习中 Actor-Critic 架构的智能体元进化

    arXiv:2606.26327v1 Announce Type: cross Abstract: In actor-critic reinforcement learning, network architectures are typically manually designed. Automating this design is challenging because each candidate must be trained before evaluation, and the design space is open-ended. To …

  222. Hugging Face Daily Papers TIER_1 English(EN) ·

    NormGuard:流匹配强化学习中的保奖励范数约束

    Reinforcement learning post-training degrades perceptual quality in flow-based generators through velocity norm inflation, which requires training-time intervention rather than inference-time corrections to maintain both reward alignment and image quality.

  223. arXiv cs.LG TIER_1 English(EN) · Yuxiong He ·

    无需真实答案的强化学习可改进LLM

    Reinforcement learning with verifiable rewards (RLVR) for training LLMs typically rely on ground-truth answers to assign rewards, limiting their applicability to tasks where the ground-truth solution is unknown. We introduce a \textbf{R}anking-\textbf{i}nduced \textbf{VER}ifiable…

  224. arXiv cs.AI TIER_1 English(EN) · Chao Yu ·

    大型语言模型智能体的经验规则与策略联合学习

    For LLM agents in multi-step interactive environments, a key challenge is to make effective use of accumulated interaction experience. Existing work has typically separated two uses of such experience: keeping it outside the model as natural-language rules for later prompting, or…

  225. arXiv cs.LG TIER_1 English(EN) · Vincent François-Lavet ·

    状态表示在深度强化学习中的重要性:应用于能源交易

    Energy trading decisions depend not only on current market prices, but also on expected future market conditions, and operational constraints. This makes the state representation given to a reinforcement learning agent an important design choice. We study this in HydroDam, a pump…

  226. arXiv cs.LG TIER_1 English(EN) · Minxian Xu ·

    RolloutPipe:在分离式策略内强化学习大模型中实现重叠流水线式部署与训练

    Large language model (LLM) post-training for reasoning increasingly relies on reinforcement learning with verifiable rewards (RLVR), where models learn from ground-truth feedback on mathematical, logical, and scientific tasks. To enable flexible resource allocation and support he…

  227. arXiv cs.LG TIER_1 English(EN) · Daoyuan Chen ·

    GEOALIGN:用于鲁棒LLM强化学习的几何展开式策展

    Online reinforcement learning is widely used to align large language models (LLMs) with reward signals, yet training can be unstable under noisy or misspecified rewards. We identify a failure mode we call directional inconsistency: within a batch, a small set of high-reward rollo…

  228. arXiv cs.CL TIER_1 English(EN) · Jianhua Tao ·

    OPID:用于智能体强化学习的策略内技能蒸馏

    Outcome-based reinforcement learning provides a stable optimization backbone for language agents, but its sparse trajectory-level rewards provide little guidance on which intermediate decisions should be reinforced or suppressed. On-policy self-distillation offers dense token-lev…

  229. arXiv cs.CL TIER_1 English(EN) · Renfei Zhang, Manasa Kaniselvan, Rylan Schaeffer abd Niloofar Mireshghallah ·

    强化学习改进LLM参数化知识的遍历

    arXiv:2511.05933v2 Announce Type: replace Abstract: Reinforcement learning (RL) is often credited with improving language model reasoning at the expense of knowledge. We challenge this narrative by showing that reasoning models consistently outperform their instruction-tuned vers…

  230. arXiv cs.LG TIER_1 English(EN) · Yuexin Bian, Jie Feng, Tao Wang, Yijiang Li, Sicun Gao, Yuanyuan Shi ·

    RN-D:用于在线策略强化学习的离散化分类Actor

    arXiv:2601.23075v2 Announce Type: replace Abstract: On-policy Reinforcement Learning (RL) remains a dominant paradigm for continuous control, yet standard implementations rely on Gaussian actors and relatively shallow MLP policies, often leading to brittle optimization when gradi…

  231. arXiv cs.LG TIER_1 English(EN) · Caleb Ju, Guanghui Lan ·

    在线强化学习的自动探索

    arXiv:2512.06244v2 Announce Type: replace Abstract: The exploration-exploitation dilemma in reinforcement learning (RL) is a fundamental challenge to efficient RL algorithms. Existing algorithms for finite state and action discounted RL problems address this by assuming sufficien…

  232. arXiv cs.LG TIER_1 English(EN) · Thiago Thomas, Gabriel de Oliveira Ramos, Felipe Meneguzzi ·

    基于团队和目标条件强化学习及因子分解分支定界的多元智能体目标识别

    arXiv:2606.25978v1 Announce Type: cross Abstract: Multi-agent goal recognition asks an observer to jointly infer which agents act together and what each team is trying to achieve, so the hypothesis space grows combinatorially with the number of team partitions and goals per team.…

  233. arXiv cs.LG TIER_1 English(EN) · Peng Xu, Sijia Chen, Junzhuo Li, Xuming Hu ·

    LLM智能体强化学习的语义一致性策略优化

    arXiv:2606.25852v1 Announce Type: new Abstract: Group-based reinforcement learning effectively post-trains LLM agents for long-horizon, sparse-reward tasks by deriving step-level credit from trajectory outcomes. However, this ties a step's credit to its rollout's final outcome: s…

  234. arXiv cs.LG TIER_1 English(EN) · Samuel Valland Lyngset, Tor Viljen Raanaas, Gard Sveipe, Eirik M{\o}ller Nilsen, Jim Torresen, Kai Olav Ellefsen, Tobias L{\o}mo ·

    强化学习中低秩适应的高效策略库

    arXiv:2606.25700v1 Announce Type: new Abstract: When fine-tuning Large Language Models (LLMs), there has been success in minimizing both memory usage and computation with Parameter-Efficient Fine-Tuning (PEFT), like Low Rank Adaptation (LoRA). In this article, we have explored wh…

  235. arXiv cs.LG TIER_1 English(EN) · Guozheng Ma, Lu Li, Zilin Wang, Pierre-Luc Bacon, Dacheng Tao ·

    超越“一刀切”:基于诊断的在线强化学习与离线先验

    arXiv:2606.25527v1 Announce Type: new Abstract: Online reinforcement learning (RL) agents increasingly depend on knowledge acquired offline to achieve practical efficiency. Originally studied in offline-to-online RL, this paradigm now spans foundation model post-training and embo…

  236. arXiv cs.LG TIER_1 English(EN) · Bang Giang Le, Viet Cuong Ta ·

    合作多智能体强化学习中具有独立参与者和顺序更新的低方差信任区域优化

    arXiv:2606.25526v1 Announce Type: new Abstract: Cooperative multi-agent reinforcement learning assumes each agent shares the same reward function and can be trained effectively using the Trust Region framework of single-agent. Instead of relying on other agents' actions, the inde…

  237. arXiv cs.LG TIER_1 English(EN) · Yivan Zhang, Ziyan Luo, Manuel Baltieri ·

    强化学习中状态抽象的组合行为语义

    arXiv:2606.25357v1 Announce Type: new Abstract: State abstraction plays a key role in scaling reinforcement learning to complex but structured systems. In studying such systems, a wide range of behavioral structures have been studied in reinforcement learning, including value fun…

  238. arXiv cs.LG TIER_1 English(EN) · Zhengzhu Liu, Zeming Gao, Haoyuan Qin, Jiawei Hu, Junhao Wu, Miao Zhu, Haipeng Zhang, Chennan Ma, Siqi Shen, Cheng Wang ·

    停滞的神经元:理解多智能体强化学习价值分解方法中的可塑性损失

    arXiv:2606.25335v1 Announce Type: new Abstract: Multi-Agent Reinforcement Learning (MARL) value factorization methods can suffer from a loss of plasticity, gradually failing to adapt when transferring to new task instances. We trace this issue to stagnant neurons, units whose gra…

  239. arXiv cs.LG TIER_1 English(EN) · Animesh Animesh, Satheesh K Perepu, Kaushik Dey ·

    GCT-MARL:基于图的对比式迁移,用于样本高效的合作多智能体强化学习

    arXiv:2606.25073v1 Announce Type: new Abstract: In cooperative multi-agent reinforcement learning (MARL), from a deployment perspective, it is challenging and expensive to train agents from scratch for each new environment or task. In this work, we propose GCT-MARL, a transfer le…

  240. arXiv cs.LG TIER_1 English(EN) · Ankur Naskar, Swetha Ganesh, Vaneet Aggarwal ·

    偏差控制的原始对偶自然Actor-Critic:约束多目标平均奖励强化学习的最优率

    arXiv:2606.25012v1 Announce Type: new Abstract: Many reinforcement learning (RL) problems in the infinite-horizon average-reward setting require optimizing multiple conflicting objectives while satisfying multiple safety constraints. A common approach is concave scalarization, wh…

  241. arXiv cs.LG TIER_1 English(EN) · Wenyang Hu, Junxiang Jia, Zhen Shu, Daniel Dahlmeier, See-Kiong Ng, Bryan Kian Hsiang Low ·

    ExTra:语言模型强化学习的探索性轨迹优化

    arXiv:2606.24994v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) for language-model reasoning can fail at both extremes of task difficulty: easy prompts often produce all-correct, low-diversity rollout groups with little gradient signal, while…

  242. arXiv cs.LG TIER_1 English(EN) · Thibaut Kulak ·

    迈向量子多任务强化学习与大型决策模型

    arXiv:2606.24962v1 Announce Type: new Abstract: Recent progress in large-scale sequence modeling has shown that a single model can learn useful representations across highly diverse data distributions. Inspired by these advances, we investigate whether a unified transformer polic…

  243. arXiv cs.LG TIER_1 English(EN) · Haoyuan Deng, Yihong Zhou, Thomas Morstyn, Yi Wang ·

    面向分布式能源资源协调的监督强化学习

    arXiv:2606.24947v1 Announce Type: new Abstract: The increasing integration of distributed energy resources (DERs) is crucial for power system decarbonization, yet unlocking DERs' flexibility is challenged by their inherent uncertainties and modelling complexity. As traditional op…

  244. arXiv cs.CL TIER_1 English(EN) · Yupu Hao, Zhuoran Jin, Huanxuan Liao, Kang Liu, Jun Zhao ·

    多步工具使用强化学习为何会崩溃以及监督信号如何修复它

    arXiv:2606.26027v1 Announce Type: new Abstract: Tool use enables large language models (LLMs) to perform complex tasks, and recent agentic reinforcement learning (RL) methods show promise for enhancing model capabilities. However, RL alone often leads to instability or limited ga…

  245. arXiv cs.CL TIER_1 English(EN) · Wenxuan Jiang, Zining Fan, Zijian Zhang, Xuecheng Wu, Hongming Tan, Haoyang Dai, Xiaoyu Li, Xuezhi Cao, Ninghao Liu ·

    OPERA:通过基于客观困惑度的强化学习实现开放式推理的对齐

    arXiv:2606.25757v1 Announce Type: new Abstract: Reinforcement Learning (RL) has enabled LLMs to excel in objective reasoning tasks such as mathematics and code generation. However, applying RL to open-ended tasks, such as creative writing, remains challenging because LLM-as-a-jud…

  246. arXiv cs.CL TIER_1 English(EN) · Hanyang Wang, Weijieying Ren, Yuxiang Zhang, Ding Cao, Zhizhao Zeng, Ke Zeng, Tianxiang Zhao ·

    BiPACE:基于双模拟引导策略优化和动作反事实估计的LLM智能体

    arXiv:2606.25556v1 Announce Type: new Abstract: Stepwise group-based RL is an attractive way to train long-horizon LLM agents without a learned critic: it reuses multiple sampled rollouts to estimate local advantages. Its weakness is less visible but more fundamental: every group…

  247. arXiv cs.LG TIER_1 English(EN) · Helai Huang ·

    通过自适应奖励塑形和策略比重加权策略实现样本高效的迁移强化学习

    Transfer learning improves policy learning efficiency by reusing knowledge from source tasks, providing a feasible paradigm for safe and efficient autonomous highway lane changing decision-making. Existing methods frequently encounter transfer mismatch induced by distribution shi…

  248. Hugging Face Daily Papers TIER_1 English(EN) ·

    Mean-Field PhiBE:从离散时间数据中进行连续时间均值场强化学习

    This paper addresses model-free continuous-time mean-field control in a setting where the population dynamics evolve continuously according to an unknown McKean-Vlasov stochastic differential equation, while only discrete-time transition data are available. In the model-based for…

  249. arXiv cs.LG TIER_1 English(EN) · Yuhua Zhu ·

    Mean-Field PhiBE:从离散时间数据中进行连续时间均值场强化学习

    This paper addresses model-free continuous-time mean-field control in a setting where the population dynamics evolve continuously according to an unknown McKean-Vlasov stochastic differential equation, while only discrete-time transition data are available. In the model-based for…

  250. Hugging Face Daily Papers TIER_1 English(EN) ·

    OPID:面向智能体强化学习的策略内技能蒸馏

    On-policy skill distillation framework extracts dense hindsight supervision from completed trajectories to improve language agent training efficiency and performance.

  251. arXiv cs.CL TIER_1 English(EN) · Jun Zhao ·

    为何多步工具使用强化学习会崩溃以及监督信号如何解决它

    Tool use enables large language models (LLMs) to perform complex tasks, and recent agentic reinforcement learning (RL) methods show promise for enhancing model capabilities. However, RL alone often leads to instability or limited gains in tool-use tasks. In our experiments, some …

  252. Hugging Face Daily Papers TIER_1 English(EN) ·

    基于团队和目标条件强化学习及因子化分支定界的多元智能体目标识别

    Multi-agent goal recognition asks an observer to jointly infer which agents act together and what each team is trying to achieve, so the hypothesis space grows combinatorially with the number of team partitions and goals per team. Real applications such as drone surveillance and …

  253. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Felipe Meneguzzi ·

    基于团队和目标条件强化学习及因子化分支定界的多元智能体目标识别

    Multi-agent goal recognition asks an observer to jointly infer which agents act together and what each team is trying to achieve, so the hypothesis space grows combinatorially with the number of team partitions and goals per team. Real applications such as drone surveillance and …

  254. Hugging Face Daily Papers TIER_1 English(EN) ·

    LLM智能体强化学习的语义一致性策略优化

    Group-based reinforcement learning effectively post-trains LLM agents for long-horizon, sparse-reward tasks by deriving step-level credit from trajectory outcomes. However, this ties a step's credit to its rollout's final outcome: semantically near-identical intermediate steps re…

  255. arXiv cs.AI TIER_1 English(EN) · Xuming Hu ·

    面向LLM智能体强化学习的语义一致性策略优化

    Group-based reinforcement learning effectively post-trains LLM agents for long-horizon, sparse-reward tasks by deriving step-level credit from trajectory outcomes. However, this ties a step's credit to its rollout's final outcome: semantically near-identical intermediate steps re…

  256. arXiv cs.CL TIER_1 English(EN) · Ninghao Liu ·

    OPERA:通过基于客观困惑度的强化学习实现开放式推理对齐

    Reinforcement Learning (RL) has enabled LLMs to excel in objective reasoning tasks such as mathematics and code generation. However, applying RL to open-ended tasks, such as creative writing, remains challenging because LLM-as-a-judge reward models often exhibit stylistic biases …

  257. Hugging Face Daily Papers TIER_1 English(EN) ·

    OPERA:通过基于客观困惑度的强化学习实现开放式推理对齐

    Reinforcement Learning (RL) has enabled LLMs to excel in objective reasoning tasks such as mathematics and code generation. However, applying RL to open-ended tasks, such as creative writing, remains challenging because LLM-as-a-judge reward models often exhibit stylistic biases …

  258. arXiv cs.LG TIER_1 English(EN) · Tobias Lømo ·

    强化学习中低秩适应的高效策略库

    When fine-tuning Large Language Models (LLMs), there has been success in minimizing both memory usage and computation with Parameter-Efficient Fine-Tuning (PEFT), like Low Rank Adaptation (LoRA). In this article, we have explored whether this approach is transferable to the world…

  259. arXiv cs.AI TIER_1 English(EN) · Tianxiang Zhao ·

    BiPACE:基于双模拟引导的策略优化与行动反事实估计,用于LLM智能体

    Stepwise group-based RL is an attractive way to train long-horizon LLM agents without a learned critic: it reuses multiple sampled rollouts to estimate local advantages. Its weakness is less visible but more fundamental: every group-relative estimator assumes that the steps it co…

  260. arXiv cs.LG TIER_1 English(EN) · Dacheng Tao ·

    超越“一刀切”:基于诊断的在线强化学习与离线先验

    Online reinforcement learning (RL) agents increasingly depend on knowledge acquired offline to achieve practical efficiency. Originally studied in offline-to-online RL, this paradigm now spans foundation model post-training and embodied intelligence, with prior types expanding fr…

  261. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Viet Cuong Ta ·

    合作多智能体强化学习中具有独立参与者和顺序更新的低方差信任区域优化

    Cooperative multi-agent reinforcement learning assumes each agent shares the same reward function and can be trained effectively using the Trust Region framework of single-agent. Instead of relying on other agents' actions, the independent actors setting considers each agent to a…

  262. Hugging Face Daily Papers TIER_1 English(EN) ·

    合作多智能体强化学习中具有独立参与者和顺序更新的低方差信任区域优化

    Cooperative multi-agent reinforcement learning assumes each agent shares the same reward function and can be trained effectively using the Trust Region framework of single-agent. Instead of relying on other agents' actions, the independent actors setting considers each agent to a…

  263. arXiv cs.AI TIER_1 English(EN) · Elias Bareinboim, Junzhe Zhang, Sanghack Lee ·

    因果强化学习简介

    arXiv:2606.24160v1 Announce Type: new Abstract: Causal inference provides a set of principles and tools that allow one to combine data and knowledge about an environment to reason with questions of counterfactual nature, i.e., what would have happened had reality been different, …

  264. arXiv cs.AI TIER_1 English(EN) · Akshay V. Jagadeesh, Rahul K. Arora, Khaled Saab, Ali Malik, Mikhail Trofimov, Foivos Tsimpourlas, Johannes Heidecke, Karan Singhal ·

    强化学习助力实现广泛且持久有益的模型

    arXiv:2606.24014v1 Announce Type: new Abstract: As AI systems are deployed across increasingly diverse and high-stakes settings, model alignment must generalize beyond the tasks and domains seen during training. This is especially important for reinforcement learning (RL), which …

  265. arXiv cs.AI TIER_1 English(EN) · Tianyuan Shi, Canbin Huang, Bei Li, Xin Chen, Xiaojun Quan, Jingang Wang, Qifan Wang ·

    超越轨迹模仿:面向大语言模型推理的策略引导策略优化

    arXiv:2606.24064v1 Announce Type: new Abstract: Distilling reasoning capabilities from strong to weak language models typically involves imitating specific solution trajectories, effectively transferring what to answer rather than how to reason. This trajectory-level imitation en…

  266. arXiv cs.LG TIER_1 English(EN) · Luca Viano, Till Freihaut, Emanuele Nevali, Volkan Cevher, Matthieu Geist, Giorgia Ramponi ·

    基于函数逼近的多智能体模仿学习:线性马尔可夫博弈及其他

    arXiv:2602.22810v2 Announce Type: replace Abstract: In this work, we present the first theoretical analysis of multi-agent imitation learning (MAIL) in linear Markov games where both the transition dynamics and each agent's reward function are linear in some given features. We de…

  267. arXiv cs.AI TIER_1 English(EN) · Anurag Akula, Satheesh K. Perepu, Abhishek Sarkar, Kaushik Dey ·

    ASALT:多智能体强化学习中横向迁移的自适应状态对齐

    arXiv:2606.24601v1 Announce Type: new Abstract: Multi-agent reinforcement learning (MARL) addresses the problem of training multiple agents that pursue collaborative, competitive, or mixed objectives. Prior work has investigated transfer learning between source and target domains…

  268. arXiv cs.AI TIER_1 English(EN) · Bingnan Xiao, Chenhao Yang, Wei Ni, Xin Wang, Tony Q. S. Quek ·

    用于策略驱动物理层系统的双层长期优化的Agentic AI

    arXiv:2606.24416v1 Announce Type: new Abstract: Network operators' changing policies, service requirements, and stringent real-time constraints render existing methods designed with fixed objectives and constraints ineffective. This paper presents Agentic long-term performance op…

  269. arXiv cs.LG TIER_1 English(EN) · Manuel Baltieri ·

    强化学习中状态抽象的组合行为语义

    State abstraction plays a key role in scaling reinforcement learning to complex but structured systems. In studying such systems, a wide range of behavioral structures have been studied in reinforcement learning, including value functions, invariants, bisimulation relations, and …

  270. Hugging Face Daily Papers TIER_1 English(EN) ·

    Stagnant Neuron: Towards Understanding the Plasticity Loss in Multi-Agent Reinforcement Learning Value Factorization Methods

    Multi-Agent Reinforcement Learning (MARL) value factorization methods can suffer from a loss of plasticity, gradually failing to adapt when transferring to new task instances. We trace this issue to stagnant neurons, units whose gradient updates become negligibly small relative t…

  271. arXiv cs.LG TIER_1 English(EN) · Cheng Wang ·

    停滞的神经元:理解多智能体强化学习价值分解方法中的可塑性损失

    Multi-Agent Reinforcement Learning (MARL) value factorization methods can suffer from a loss of plasticity, gradually failing to adapt when transferring to new task instances. We trace this issue to stagnant neurons, units whose gradient updates become negligibly small relative t…

  272. Hugging Face Daily Papers TIER_1 English(EN) ·

    多步工具使用强化学习为何会崩溃以及监督信号如何修复它

    Research investigates how different supervisory signals and training strategies improve the stability and performance of large language models in tool-use tasks, addressing issues like catastrophic collapse and format sensitivity through interleaved supervised fine-tuning and rei…

  273. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Kaushik Dey ·

    GCT-MARL:基于图的对比式迁移,用于样本高效的合作多智能体强化学习

    In cooperative multi-agent reinforcement learning (MARL), from a deployment perspective, it is challenging and expensive to train agents from scratch for each new environment or task. In this work, we propose GCT-MARL, a transfer learning framework that builds on the multi-view g…

  274. arXiv cs.AI TIER_1 English(EN) · Kaushik Dey ·

    ASALT:多智能体强化学习中横向迁移的自适应状态对齐

    Multi-agent reinforcement learning (MARL) addresses the problem of training multiple agents that pursue collaborative, competitive, or mixed objectives. Prior work has investigated transfer learning between source and target domains in MARL; however, the majority of existing appr…

  275. arXiv cs.AI TIER_1 English(EN) · Tony Q. S. Quek ·

    用于策略驱动物理层系统的双层长期优化的Agentic AI

    Network operators' changing policies, service requirements, and stringent real-time constraints render existing methods designed with fixed objectives and constraints ineffective. This paper presents Agentic long-term performance optimization (Agentic-LTPO), a nested bilevel opti…

  276. Hugging Face Daily Papers TIER_1 English(EN) ·

    面向分布式能源资源协调的监督强化学习

    The increasing integration of distributed energy resources (DERs) is crucial for power system decarbonization, yet unlocking DERs' flexibility is challenged by their inherent uncertainties and modelling complexity. As traditional optimization methods struggle with such uncertaint…

  277. Hugging Face Daily Papers TIER_1 English(EN) ·

    因果强化学习简介

    Causal inference provides a set of principles and tools that allow one to combine data and knowledge about an environment to reason with questions of counterfactual nature, i.e., what would have happened had reality been different, even when no data of this unrealized reality is …

  278. arXiv cs.CL TIER_1 English(EN) · Karan Singhal ·

    强化学习助力实现广泛且持久有益的模型

    As AI systems are deployed across increasingly diverse and high-stakes settings, model alignment must generalize beyond the tasks and domains seen during training. This is especially important for reinforcement learning (RL), which can introduce unexpected misalignment through re…

  279. arXiv cs.CL TIER_1 English(EN) · Qi Zhang ·

    面向长时域智能体强化学习的群图策略优化

    Group-based Reinforcement Learning (RL) has significantly enhanced Large Language Models (LLMs) in agentic scenarios. To achieve finer-grained policy updates, recent agentic RL frameworks have shifted from trajectory-level to step-level training. However, long-horizon agentic RL …

  280. arXiv cs.CL TIER_1 English(EN) · SingGuard Team ·

    SingGuard:一种具有动态推理能力的策略自适应多模态大模型防护栏

    Vision-language models (VLMs) are increasingly deployed in consumer, medical, financial, and enterprise applications. This broad deployment expands the safety surface: risks can arise from multimodal question answering, assistant responses, and cross-modal composition, while mode…

  281. Hugging Face Daily Papers TIER_1 English(EN) ·

    Select-to-Act:通过自适应语言指导实现分层强化学习

    Reinforcement Learning (RL) has been widely applied to sequential decision-making, yet it often suffers from poor sample efficiency due to costly interactions with the environment. A limited line of recent work has started exploring improving RL efficiency by leveraging external …

  282. arXiv cs.AI TIER_1 English(EN) · Yanxi Chen, Weijie Shi, Yuexiang Xie, Boyi Hu, Yaliang Li, Bolin Ding, Jingren Zhou ·

    连接点滴:通过强化学习训练具备跨域泛化能力的长期生命周期智能体LLM

    arXiv:2606.20002v1 Announce Type: cross Abstract: This work presents a general framework for training large language models (LLMs) to "Connect the Dots" (CoD), a meta-capability required by long-lifecycle agents: as an LLM-based AI agent gets deployed in an environment, it solves…

  283. arXiv cs.AI TIER_1 English(EN) · Jingren Zhou ·

    连接点滴:通过强化学习训练具备跨领域泛化能力的长期生命周期智能体LLM

    This work presents a general framework for training large language models (LLMs) to "Connect the Dots" (CoD), a meta-capability required by long-lifecycle agents: as an LLM-based AI agent gets deployed in an environment, it solves a long sequence of tasks while continuously explo…

  284. Hugging Face Daily Papers TIER_1 English(EN) ·

    连接点滴:通过强化学习训练具备跨领域泛化能力的长期生命周期智能体LLM

    Large language models can be trained through reinforcement learning to develop a meta-capability enabling continuous learning and adaptation across long sequences of tasks in dynamic environments.

  285. arXiv stat.ML TIER_1 English(EN) · Sara Rajaram, R. James Cotton, Fabian H. Sinz ·

    相似性作为奖励对齐:基于偏好的鲁棒且通用的强化学习

    arXiv:2506.12529v2 Announce Type: replace-cross Abstract: Preference-based Reinforcement Learning (PbRL) entails a variety of approaches for aligning models with human intent to alleviate the burden of reward engineering. However, most previous PbRL work has not investigated the …

  286. arXiv stat.ML TIER_1 English(EN) · Hari Prasad ·

    审计分布强化学习的风险声明

    arXiv:2607.11607v1 Announce Type: cross Abstract: Distributional reinforcement learning agents learn full return distributions that are increasingly read at face value: for interpretability, risk-sensitive control, and safety monitoring. We ask a question theory anticipates but t…

  287. arXiv stat.ML TIER_1 English(EN) · Wen-Ting Wang ·

    动态费用下闭环DEX模拟器中的执行强化学习

    arXiv:2607.10960v1 Announce Type: cross Abstract: Trader-facing dynamic fees are increasingly proposed for automated market makers (AMMs), but historical data do not identify how order flow would respond: trader-facing fees do not vary, trader types are latent, and a replayed tap…

  288. arXiv stat.ML TIER_1 English(EN) · Asim H. Gazi, Yongyi Guo, Daiqi Gao, Ziping Xu, Kelly W. Zhang, Susan A. Murphy ·

    现实世界中的强化学习:统计挑战与未来方向调查

    arXiv:2601.15353v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has achieved remarkable success in real-world decision-making across diverse domains, including gaming, robotics, online advertising, public health, and natural language processing. Despite thes…

  289. arXiv stat.ML TIER_1 English(EN) · Yang Xu, Swetha Ganesh, Vaneet Aggarwal ·

    高效Q学习和Actor-Critic方法用于鲁棒平均奖励强化学习

    arXiv:2506.07040v4 Announce Type: replace-cross Abstract: We study model-free methods for distributionally robust infinite-horizon average-reward Markov decision processes (MDPs). We present non-asymptotic convergence analyses of Q-learning and actor-critic algorithms for robust …

  290. arXiv stat.ML TIER_1 English(EN) · Miao Lu, Han Zhong, Tong Zhang, Jose Blanchet ·

    分布鲁棒强化学习与交互式数据收集:基础硬度和近乎最优算法

    arXiv:2404.03578v3 Announce Type: replace-cross Abstract: The sim-to-real gap, which represents the disparity between training and testing environments, poses a significant challenge in reinforcement learning (RL). A promising approach to addressing this challenge is distribution…

  291. arXiv stat.ML TIER_1 English(EN) · Hari Prasad ·

    审计分布强化学习的风险声明

    Distributional reinforcement learning agents learn full return distributions that are increasingly read at face value: for interpretability, risk-sensitive control, and safety monitoring. We ask a question theory anticipates but that has not been measured directly: are the risk c…

  292. arXiv stat.ML TIER_1 English(EN) · Wen-Ting Wang ·

    动态费用下闭环DEX模拟器中的执行强化学习

    Trader-facing dynamic fees are increasingly proposed for automated market makers (AMMs), but historical data do not identify how order flow would respond: trader-facing fees do not vary, trader types are latent, and a replayed tape is not a sequential decision environment. We the…

  293. arXiv stat.ML TIER_1 English(EN) · Zijie Cheng, Yang Peng, Zhihua Zhang ·

    Quantile Distributional Reinforcement Learning 的统计效率与推理

    arXiv:2607.08444v1 Announce Type: new Abstract: In this paper, we study quantile-based distributional reinforcement learning from the perspective of statistical efficiency. We focus on distributional policy evaluation, whose goal is to characterize the return distribution, namely…

  294. arXiv stat.ML TIER_1 English(EN) · Zhihua Zhang ·

    Quantile Distributional Reinforcement Learning 的统计效率和推理

    In this paper, we study quantile-based distributional reinforcement learning from the perspective of statistical efficiency. We focus on distributional policy evaluation, whose goal is to characterize the return distribution, namely the distribution of discounted cumulative rewar…

  295. arXiv stat.ML TIER_1 English(EN) · Zhihua Zhang ·

    Quantile Distributional Reinforcement Learning 的统计效率与推理

    In this paper, we study quantile-based distributional reinforcement learning from the perspective of statistical efficiency. We focus on distributional policy evaluation, whose goal is to characterize the return distribution, namely the distribution of discounted cumulative rewar…

  296. arXiv stat.ML TIER_1 English(EN) · Nikita Yudin ·

    强化学习的数学方法

    Reinforcement learning (RL) is increasingly grounded in tools from probability, optimization, and operator theory. This survey organizes the mathematical structures that underpin the design and analysis of modern algorithms in RL. We begin from Markov decision processes (MDPs) an…

  297. arXiv cs.CV TIER_1 English(EN) · Dengyang Jiang, Dongyang Liu, Zanyi Wang, Qilong Wu, Liuzhuozheng Li, Hengzhuang Li, Xin Jin, David Liu, Changsheng Lu, Zhen Li, Bo Zhang, Mengmeng Wang, Steven Hoi, Peng Gao, Harry Yang ·

    分布匹配蒸馏遇上强化学习

    arXiv:2511.13649v5 Announce Type: replace Abstract: Distribution Matching Distillation (DMD) facilitates efficient inference by distilling multi-step diffusion models into few-step variants. Concurrently, Reinforcement Learning (RL) has emerged as a vital tool for aligning genera…

  298. arXiv stat.ML TIER_1 English(EN) · Stefano Masini, Cecilia Viscardi, Michela Baccini ·

    LF-IBIS 全贝叶斯强化学习

    arXiv:2607.01741v1 Announce Type: new Abstract: Reinforcement Learning (RL) is a sequential decision-making framework in which an agent learns optimal policies through interaction with an environment by maximizing cumulative rewards. Among RL methods, Bayesian Reinforcement Learn…

  299. arXiv stat.ML TIER_1 English(EN) · Michela Baccini ·

    LF-IBIS 全贝叶斯强化学习

    Reinforcement Learning (RL) is a sequential decision-making framework in which an agent learns optimal policies through interaction with an environment by maximizing cumulative rewards. Among RL methods, Bayesian Reinforcement Learning (BRL) addresses common practical challenges …

  300. arXiv stat.ML TIER_1 English(EN) · Yingru Li, Jiacai Liu, Jiawei Xu, Yuxuan Tong, Ziniu Li, Qian Liu, Baoxiang Wang ·

    用于长视野大语言模型强化学习的信任区域掩码

    arXiv:2512.23075v5 Announce Type: replace-cross Abstract: Policy gradient methods for Large Language Models optimize a policy $\pi_\theta$ via a surrogate objective computed from samples of a rollout policy $\pi_{\text{roll}}$. However, modern LLM-RL pipelines suffer from unavoid…

  301. Towards AI TIER_1 English(EN) · Fousseyni Sangaré ·

    强化学习:价值迭代在机器人导航任务中的应用

    <p>Hello everyone 😀 Welcome to this blog post !</p><p>Today we are going to see a real world application of Dynamic Programming in a Robotics Task. In a previous <a href="https://medium.com/@fousseyni.phd/rl-in-discrete-world-dynamic-programming-part2-generalized-policy-iteration…

  302. Towards AI TIER_1 English(EN) · Michael Ariaga ·

    置信感知强化学习:在动态环境中推进大型语言模型

    <h4>Building Large Language Model Predictive Confidence to Navigate Uncertainty with Resiliency and Conviction</h4><p>Environments are in constant change as the physical world and contextual signals evolve to reflect new meaning or redefine ground truth. Large language models (LL…

  303. r/LocalLLaMA TIER_1 English(EN) · /u/de4dee ·

    [2607.07508] Agentic Reinforcement Learning 的单次推广异步优化

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1uw7rm8/260707508_singlerollout_asynchronous_optimization/"> <img alt="[2607.07508] Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning" src="https://external-preview.redd.it/q3evP6JeDp…

  304. Mastodon — mastodon.social TIER_1 English(EN) · strike007 ·

    超越启发式安全,非空泛泛化界限对于强化学习中的形式化验证至关重要。如果我们无法在数学上保证

    Moving beyond heuristic safety, non-vacuous generalization bounds are critical for formal verification in reinforcement learning. If we cannot mathematically guarantee that an agent stays within its policy constraints, deploying it in high-stakes environments is reckless. # AI # …