DIA:用于扩散策略优化的去噪中间优势
DIA: Denoising Intermediate Advantage for Diffusion Policy Optimization
另 2 家 · 查看机构详情
- University of Toronto(多伦多大学)
- Vector Institute for Artificial Intelligence(向量人工智能研究所)
- Acceleration Consortium, University of Toronto(多伦多大学加速联盟)
- Canadian Institute for Advanced Research (CIFAR)(加拿大高等研究院(CIFAR))
- NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对扩散策略微调中忽略中间去噪步骤贡献的问题,提出DIA方法,通过学习部分去噪动作价值函数构建去噪级优势,结合PPO环境优势,在多个基准上提升最终性能与策略探索能力。
中文摘要 AI 辅助
基于扩散的机器人策略已广泛应用于机器人操作领域,通常通过行为克隆进行训练。然而,仅从演示中训练的策略受限于可用数据的质量和覆盖范围。强化学习可以通过交互进一步提升这些预训练策略的性能。一种常见的方法是使用策略梯度方法,将扩散策略微调构建为外部环境MDP与内部去噪MDP的组合。然而,现有方法通常将相同的环境级回报分配给用于构建动作块的所有去噪步骤,而不区分哪些中间决策对最终回报贡献最大。我们提出了去噪中间优势(DIA),一种策略梯度方法,该方法学习部分去噪动作上的价值函数,并利用该函数为生成过程的每个步骤构建去噪级优势。DIA将此内部信用信号与标准的环境级PPO优势相结合,在整个去噪链中提供状态相关的信用分配。在Robomimic、FurnitureBench、Franka Kitchen和D3IL上,DIA始终优于现有的扩散策略微调方法,提升了最终性能。除了最终奖励外,DIA能更高效地达到成功状态,并能更远地偏离预训练行为分布,从而发现基线方法无法达到的更有效、更高效的任务级策略和子任务序列。
英文摘要
Diffusion-based robot policies have become widely used in robotic manipulation, where they are typically trained with behavior cloning. However, policies trained purely from demonstrations are limited by the quality and coverage of the available data. Reinforcement learning can further improve the performance of these pretrained policies through interaction. A common approach is to use policy-gradient methods that formulate diffusion-policy fine-tuning as an outer environment MDP together with an inner denoising MDP. However, existing methods typically assign the same environment-level credit to all denoising steps used to construct an action chunk, without distinguishing which intermediate decisions contributed most to the final return. We introduce Denoising Intermediate Advantage (DIA), a policy-gradient method that learns a value function over partially denoised actions and uses it to construct a denoising level advantage for each step of the generative process. DIA combines this inner credit signal with the standard environment-level PPO advantage, providing state-dependent credit throughout the denoising chain. Across Robomimic, FurnitureBench, Franka Kitchen, and D3IL, DIA consistently improves final performance over existing diffusion-policy fine-tuning methods. Beyond final reward, DIA reaches successful states more efficiently and can shift farther from the pretrained behavior distribution, enabling it to discover more effective and efficient task-level strategies and subtask sequences that baseline methods fail to reach.