D2PPO: Diffusion Policy Policy Optimization with Dispersive Loss
D2PPO:带有分散损失的扩散策略策略优化
机构 * School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院) ; Guangdong Key Laboratory of Big Data Analysis and Processing(广东省大数据分析与处理重点实验室)
AI总结 本文提出D2PPO,通过引入分散损失正则化来解决扩散策略中表示崩溃问题,提升复杂机器人操控任务的性能,实验表明其在预训练和微调中均取得显著改进。
Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, 40(22): 18891-18899, 2026