研究人员正在开发新方法来提高扩散模型在强化学习和规划任务中的鲁棒性和安全性。一种方法是鲁棒正则化策略迭代(RRPI),它通过针对最坏情况动力学进行优化来解决转移不确定性,并在 D4RL 基准测试中表现出强劲的性能。另一组论文介绍了 Kolmogorov Regression 和 DiRecT 等技术,通过提高轨迹规律性来增强扩散策略,从而实现确定性故障检测,并在推理过程中强制执行安全约束,而不会过度约束采样过程。这些进展旨在使扩散模型在复杂的、长期的任务和安全关键型应用中更加可靠。
AI
arXiv:2602.05533v3 Announce Type: replace Abstract: We study conditional generation in diffusion models under hard constraints, where generated samples must satisfy prescribed events with probability one. Such constraints arise naturally in safety-critical applications and in rar…
arXiv:2603.09344v3 Announce Type: replace Abstract: Offline reinforcement learning (RL) enables data-efficient and safe policy learning without online exploration, but its performance often degrades under distribution shift. The learned policy may visit out-of-distribution state-…
Finite-dimensional (FD) diffusion policies exhibit temporal drift owing to discretization artifacts that degrade long-horizon performance (when deployed on physical systems). We introduce a backward Kolmogorov equation that lifts diffusion policies to a Cameron-Martin space -- a …
arXiv:2606.15359v1 Announce Type: new Abstract: Diffusion models have emerged as powerful tools for planning and control by learning multimodal distributions over actions and trajectories. Yet reliable inference-time safety enforcement remains a key barrier to their deployment in…
arXiv cs.AI
TIER_1English(EN)·Abhinav Agarwal, Adam Wei, Taylan Kargin, Michael Zeng, Cole Becker, Arif Kerem Dayi, Pablo Parrilo, Asuman Ozdaglar, Russ Tedrake·
arXiv:2606.16447v1 Announce Type: cross Abstract: Imitation learning has enabled highly-dexterous robotic manipulation from RGB observations. Policies trained with these methods, however, typically condition robot actions on only a short history of observations. These policies ca…
arXiv:2606.13795v1 Announce Type: new Abstract: RL post-training has become increasingly pivotal for improving diffusion policies, but existing diffusion policy-gradient methods are often unstable and cannot achieve reliable policy improvement. We identify the cause as the double…