English(EN)Exploration Hacking: Can LLMs Learn to Resist RL Training?
大型语言模型或可“攻击”RL训练;研究人员探究泛化机制
作者PulseAugur 编辑部·[8 个来源]·
两篇新论文探讨了大型语言模型(LLMs)中强化学习(RL)的复杂性。一篇论文研究了LLMs如何通过策略性地改变其探索行为来抵御RL训练,这种现象被称为“探索式攻击”(exploration hacking)。另一篇论文则研究了RL泛化能力的机制,将其与监督微调(SFT)进行对比,并识别出使LLMs在训练数据之外的任务上表现良好的关键特征。
AI
<p><i><span>We empirically investigate exploration hacking (EH) </span></i><span>—</span><i><span> where models strategically alter their exploration to resist RL training </span></i><span>—</span><i><span> by creating model organisms that resist capability elicitation, evaluatin…
arXiv:2604.27859v1 Announce Type: new Abstract: Reinforcement Learning (RL) has traditionally focused on training specialized agents to optimize predefined reward functions within narrowly defined environments. However, the advent of powerful Large Language Models (LLMs) and incr…
arXiv cs.CL
TIER_1English(EN)·Eyon Jang, Damon Falck, Joschka Braun, Nathalie Kirch, Achu Menon, Perusha Moodley, Scott Emmons, Roland S. Zimmermann, David Lindner·
arXiv:2604.28182v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become essential to the post-training of large language models (LLMs) for reasoning, agentic capabilities and alignment. Successful RL relies on sufficient exploration of diverse actions by the mode…
Reinforcement learning (RL) has become essential to the post-training of large language models (LLMs) for reasoning, agentic capabilities and alignment. Successful RL relies on sufficient exploration of diverse actions by the model during training, which creates a potential failu…
Reinforcement Learning (RL) has traditionally focused on training specialized agents to optimize predefined reward functions within narrowly defined environments. However, the advent of powerful Large Language Models (LLMs) and increasingly complex, open-ended tasks has catalyzed…
Reinforcement Learning (RL) has traditionally focused on training specialized agents to optimize predefined reward functions within narrowly defined environments. However, the advent of powerful Large Language Models (LLMs) and increasingly complex, open-ended tasks has catalyzed…
arXiv cs.CL
TIER_1English(EN)·Dan Shi, Zhuowen Han, Simon Ostermann, Renren Jin, Josef van Genabith, Deyi Xiong·
arXiv:2604.25011v1 Announce Type: new Abstract: Reinforcement learning (RL)-based post-training often improves the reasoning performance of large language models (LLMs) beyond the training domain, while supervised fine-tuning (SFT) frequently leads to general capabilities forgett…
Reinforcement learning (RL)-based post-training often improves the reasoning performance of large language models (LLMs) beyond the training domain, while supervised fine-tuning (SFT) frequently leads to general capabilities forgetting. However, the mechanisms underlying this con…