arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PAC-DP:PAC贝叶斯扩散策略学习

PAC-DP: PAC-Bayesian Diffusion Policy Learning

Mohammad Hasan Yeganegi, Dian Yu, Andrea Del Prete, Majid Khadiv, Matteo Saveriano

arXiv 2607.24296首次发表:更新:

发表机构

University of Trento; Munich Institute of Robotics and Machine Intelligence (MIRMI), Technical University of Munich (TUM)(特伦托大学; 慕尼黑工业大学慕尼黑机器人与机器智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对扩散策略在有限数据下泛化控制有限的问题,提出PAC-DP,将DP建模为贝叶斯神经网络并定义泛化界导出新训练目标。该方法理论上规范DP训练,实验显示在多基准测试中性能提升,尤其在低数据和复杂任务中效果显著。

AI 中文摘要

扩散策略(DPs)能执行复杂操作任务,但通常通过最小化去噪目标进行训练,在机器人领域常见的有限数据情况下对泛化的控制有限。本文提出PAC-DP,通过将DP建模为贝叶斯神经网络并定义PAC-贝叶斯泛化界,导出新训练目标,用后验与先验参数分布间的Kullback-Leibler散度正则化器增强标准去噪损失。理论上提供了规范DP训练的原则方法,实验表明在多机器人操作基准测试中去噪性能提高、变分负对数似然降低、成功率更高,尤其在低数据训练和复杂任务中有最大改进。

英文摘要

Diffusion Policies (DPs) are able to perform complex manipulation tasks. However, DPs are typically trained by minimizing a denoising objective, which provides limited control over generalization in the finite-data regimes common in robotics. In this letter, we propose PAC-DP, an approach that increases the performance of DPs in robotic manipulation tasks. By modeling the DP as a Bayesian neural network, and defining a PAC-Bayes generalization bound, we derive a novel training objective that augments the standard denoising loss with a Kullback-Leibler divergence regularizer between the posterior and prior parameter distributions. From the theoretical perspective, our approach provides a principled approach to regularize the training of DPs without significantly increasing the training time. From the practical point of view, experimental results demonstrate improved denoising performance, lower variational negative log-likelihood, and higher success rates across multiple robotic manipulation benchmarks. Crucially, the largest improvements are observed in low-data training regimes and complex tasks, establishing PAC-DP as a theoretically grounded framework for robot policy learning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑