发表机构
University of Trento; Munich Institute of Robotics and Machine Intelligence (MIRMI), Technical University of Munich (TUM)(特伦托大学; 慕尼黑工业大学慕尼黑机器人与机器智能研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对扩散策略在有限数据下泛化控制有限的问题,提出PAC-DP,将DP建模为贝叶斯神经网络并定义泛化界导出新训练目标。该方法理论上规范DP训练,实验显示在多基准测试中性能提升,尤其在低数据和复杂任务中效果显著。
AI 中文摘要
扩散策略(DPs)能执行复杂操作任务,但通常通过最小化去噪目标进行训练,在机器人领域常见的有限数据情况下对泛化的控制有限。本文提出PAC-DP,通过将DP建模为贝叶斯神经网络并定义PAC-贝叶斯泛化界,导出新训练目标,用后验与先验参数分布间的Kullback-Leibler散度正则化器增强标准去噪损失。理论上提供了规范DP训练的原则方法,实验表明在多机器人操作基准测试中去噪性能提高、变分负对数似然降低、成功率更高,尤其在低数据训练和复杂任务中有最大改进。
英文摘要
Diffusion Policies (DPs) are able to perform complex manipulation tasks. However, DPs are typically trained by minimizing a denoising objective, which provides limited control over generalization in the finite-data regimes common in robotics. In this letter, we propose PAC-DP, an approach that increases the performance of DPs in robotic manipulation tasks. By modeling the DP as a Bayesian neural network, and defining a PAC-Bayes generalization bound, we derive a novel training objective that augments the standard denoising loss with a Kullback-Leibler divergence regularizer between the posterior and prior parameter distributions. From the theoretical perspective, our approach provides a principled approach to regularize the training of DPs without significantly increasing the training time. From the practical point of view, experimental results demonstrate improved denoising performance, lower variational negative log-likelihood, and higher success rates across multiple robotic manipulation benchmarks. Crucially, the largest improvements are observed in low-data training regimes and complex tasks, establishing PAC-DP as a theoretically grounded framework for robot policy learning.