发表机构
University of California, Berkeley; Princeton University; The University of Texas at Austin(加州大学伯克利分校; 普林斯顿大学; 德克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对扩散策略在多模态机器人动作生成中易坍缩的问题,提出不可混溶扩散策略,通过无标签噪声分配保留清晰的动作-噪声路径,在模拟和真实任务中显著提升模态保留并保持性能。
AI 中文摘要
当扩散策略首次被提出时,人们期望它们能够恢复多模态的动作分布。然而,我们发现这一期望并不总是成立,因为即使我们保证了数据集模态的平衡以及批内精确的对称性,扩散策略也常常坍缩到单一模态。我们的分析表明,独立的动作-噪声配对通过增加扩散路径之间的混合与交叉而导致这一失败,这可能产生平均化的去噪响应并抑制特定模态的行为。这个问题在机器人规划中尤为严重,因为动作空间是稠密且低维的,显著增加了这种混合与交叉。为了缓解这一问题,我们提出了不可混溶扩散策略,这是一种无标签的训练时附加组件,适用于扩散策略,通过使用动作-噪声分配来保留相对清晰的噪声到动作的路径,而无需修改策略架构或推理过程。在跨越状态、RGB和点云观测的五个模拟和两个真实世界的人形操作任务中,我们的方法显著提高了策略对动作模态的保留,同时保持了强大的任务性能。它在三个双模态任务中将非主导模态的比例提高了6.0倍至14.6倍,并在两个四模态任务中恢复了在原始策略 rollout 中完全缺失的演示模态。这些结果表明,不可混溶扩散策略提供了一种简单而稳健的方法,用于在通用机器人学习任务中保留动作多模态性。
英文摘要
When diffusion policies were first introduced, they were expected to recover multi-modal action distributions. However, we find this expectation does not always hold, as diffusion policies often collapse to a single modality even when we guarantee the balance of dataset modalities and exact within-batch symmetry. Our analysis indicates that independent action-noise pairing contributes to this failure by increasing mixing and crossing among diffusion paths, which can produce averaged denoising responses and suppress modality-specific behavior. This issue is especially severe in robot planning, where action spaces are dense and low-dimensional, significantly increasing such mixing and crossing. To alleviate this problem, we propose Immiscible Diffusion Policy, a label-free training-time add-on to diffusion policy that uses action-noise assignment to preserve relatively distinct noise-to-action routes without modifying the policy architecture or inference procedure. Across five simulated and two real-world humanoid manipulation tasks spanning state, RGB, and point-cloud observations, our method significantly improves the policy's preservation of action modalities while maintaining strong task performance. It increases the proportion of the non-dominant modality by 6.0x-14.6x across three two-modality tasks and recovers demonstrated modalities that are entirely absent from vanilla policy rollouts on both four-modality tasks. These results demonstrate that Immiscible Diffusion Policy provides a simple yet robust approach to preserving action multi-modality in general robot learning tasks.