发表机构
Dalian University of Technology(大连理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出运动趋势引导,通过潜在表示提供前瞻性,仅增3.52%参数,在多个基准上显著提升3D扩散策略性能。
AI 中文摘要
3D扩散策略在从当前观测生成几何上合理的动作方面表现出色,但成功的操作不仅需要知道当前哪些动作可行,还需要预测交互的发展方向。现有策略大多让这种前瞻性隐式地从动作学习中涌现。我们引入了运动趋势引导(Movement Trend Guidance),一种简单而有效的方法,在不引入显式计划的情况下提供这种前瞻性。策略从短暂的观测历史中学习交互演化的紧凑潜在表示。在训练期间,稀疏的未来夹爪状态监督该表示;在推理时,仅保留该潜在表示作为面向未来的条件,与当前观测一起使用。该潜在表示为动作生成提供全局条件,而额外的门控FiLM分支仅在UNet瓶颈处使用。尽管仅比DP3增加了3.52%的参数,我们的方法保留了原始的密集动作和滚动时域公式,并在RoboTwin2.0、LIBERO-40和DexArt上持续优于DP3。在50任务RoboTwin2.0混合训练中达到62.8%对比56.1%,在LIBERO-40上达到71.93%对比37.08%,在五个真实机器人任务上达到72.0%对比49.0%。这些结果表明,扩散策略能够从了解交互的发展方向中显著受益,而无需被告知确切移动位置。
英文摘要
3D diffusion policies are strong at generating geometrically grounded actions from current observations, but successful manipulation requires not only knowing what motion is feasible now, but also anticipating where the interaction is heading. Existing policies largely leave such foresight to emerge implicitly from action learning. We introduce Movement Trend Guidance, a simple but effective way to provide this foresight without introducing an explicit plan. From a short observation history, the policy learns a compact latent representation of interaction evolution. During training, sparse future gripper states supervise this representation; at inference, only the latent is retained as future-oriented conditioning alongside the current observation. The latent provides global conditioning for action generation, while an additional gated FiLM branch is used only at the UNet bottleneck. Despite adding only 3.52% more parameters to DP3, our method preserves the original dense-action and receding-horizon formulation and consistently improves upon DP3 across RoboTwin2.0, LIBERO-40, and DexArt. It reaches 62.8% vs. 56.1% in 50-task RoboTwin2.0 mixed training, 71.93% vs. 37.08% on LIBERO-40, and 72.0% vs. 49.0% on five real-robot tasks. These results show that a diffusion policy can benefit substantially from knowing where an interaction is heading, without being told exactly where to move.