arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

iADD:提升扩散策略优化中的对齐性与多样性

iADD: Improving Alignment and Diversity in Diffusion Policy Optimization

Ashok Prasad Neupane, Saugat Adhikari, Pramish Paudel, Ajad Chhatkuli, Danda Pani Paudel

arXiv 2610.01789首次发表:更新:

发表机构

Pulchowk Campus, IOE, Tribhuvan University; NAAMII(特里布万大学工程学院普尔乔克校区; NAAMII)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出iADD方法,通过理论分析证明仅更新后段时间步有害于多样性,并基于增量式Feynman-Kac训练实现扩散策略优化中对齐性与多样性的最佳权衡,实验验证了显著性能提升。

AI 中文摘要

基于强化学习的扩散模型后训练方法,如去噪扩散策略优化(DDPO),在奖励函数下优化反向扩散过程。然而,当前的奖励优化方法以牺牲多样性和质量为代价。本文通过细致的理论思考和方法设计提供了更好的权衡。我们分析了理论框架,并从数学上证明,与先前工作中的结论相反,仅对扩散模型的后段时间步进行更新可能对多样性有害。此外,我们基于坚实的理论基础提出了一种增量式Feynman-Kac训练方法,以实现迄今最佳的对齐-多样性权衡。我们进行了大量实验,在三个不同任务中将我们的方法与相关的扩散策略优化方法进行比较,并对每个组件进行了强有力的消融研究,从而验证了在对齐性和多样性方面的显著性能提升。

英文摘要

Reinforcement learning based post training of diffusion models, such as Denoising Diffusion Policy Optimization (DDPO), optimizes a reverse diffusion process under a reward function. However, current approaches to reward optimizations do so at the cost of diversity and quality. In this paper, we provide better tradeoffs through careful theoretical considerations and method design. We analyze the theoretical framework and mathematically demonstrate that \emph{only-latter timestep} updates of diffusion model may be harmful for diversity contrary to the conclusions presented in a previous work. Additionally, we propose an incremental Feynman-Kac training based on strong theoretical foundations in order to achieve the best-yet alignment-diversity tradeoffs. We perform extensive experiments and compare our method against related diffusion policy optimization approaches in three different tasks and also provide strong ablations for each component, thus validating strong performance gains in both alignment and diversity.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑