arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15570cs.RO

DIDO:将交互中心动力学蒸馏为单步去噪的世界动作模型

DIDO: Distilling Interaction-Centric Dynamics into One-Step Denoising for World Action Models

Jing Lyu, Shuanghao Bai, Runze Xiao, Zhenyu Liao, Wenxing Tan, Zihan Tang, Ruochuan Shi, Cheng Peng, Yuheng Ji, Yihao Wang, Badong Chen, Pengwei Wang, Zhongyuan… 展开作者

Jing Lyu, Shuanghao Bai, Runze Xiao, Zhenyu Liao, Wenxing Tan, Zihan Tang, Ruochuan Shi, Cheng Peng, Yuheng Ji, Yihao Wang, Badong Chen, Pengwei Wang, Zhongyuan Wang, Xiaoguang Zhao

首次发表
浏览论文内容

中文总结 AI 辅助

针对多步视频生成模型在机器人操作中延迟高的问题,提出DIDO,通过分布匹配蒸馏和交互中心表示引导将多步动力学压缩为单步去噪,在LIBERO等基准上实现高成功率并降低延迟。

中文摘要 AI 辅助

世界动作模型(WAMs)利用视频生成模型预测机器人操作中的未来视觉动态,但迭代去噪为闭环控制引入了额外延迟。我们通过实验发现,在去噪过程中视觉内容的收敛速率不同。静态背景结构形成较早,而夹爪和被操作物体在第一步后仍保持模糊,其交互动力学仅通过后续去噪步骤才逐渐显现。因此,将多步视频模型简单截断为单步虽能保留场景结构,却会丢失对操作最为关键的交互中心动力学。为解决这一问题,我们提出DIDO,将多步视频模型的收敛动力学蒸馏到单步去噪中。DIDO结合了分布匹配蒸馏与交互中心表示引导。除了将多步生成压缩为一次前向传播外,DIDO还利用有监督的边界框视觉推理标记显式建模夹爪、被操作物体及其交互。此外,DIDO将目标物体在多个模型层中的表示与预训练DINOv3编码器的特征对齐。这种交互中心引导帮助蒸馏模型在单步中同时保留相关实体及其未来动力学,同时大幅降低推理延迟。DIDO在LIBERO上达到99.0%的平均成功率,在LIBERO-Plus上为76.6%,在RoboTwin上为92.0%,并展示了在真实机器人操作中向长时程和泛化任务的有效迁移。

英文摘要

World Action Models (WAMs) use video generation models to predict future visual dynamics for robotic manipulation, but iterative denoising introduces additional latency for closed-loop control. We empirically find that visual content converges at different rates during denoising. Static background structure forms early, whereas the gripper and manipulated object remain blurry after the first step, with their interaction dynamics emerging only through subsequent denoising. Consequently, naively truncating a multi-step video model to one step preserves scene structure but loses the interaction-centric dynamics most critical for manipulation. To address this issue, we propose DIDO, which distills the converged dynamics of a multi-step video model into a single denoising step. DIDO combines distribution matching distillation with interaction-centric representation guidance. Beyond compressing multi-step generation into one forward pass, DIDO explicitly models the gripper, manipulated object, and their interaction using supervised bounding-box visual reasoning tokens. Additionally, DIDO aligns the target object's representations across multiple model layers with features from a pretrained DINOv3 encoder. This interaction-centric guidance helps the distilled model preserve both the relevant entities and their future dynamics in a single step, while substantially reducing inference latency. DIDO achieves an average success rate of 99.0\% on LIBERO, 76.6\% on LIBERO-Plus, and 92.0\% on RoboTwin, while also demonstrating effective transfer to long-horizon and generalization tasks in real-world robotic manipulation.

发表机构

  • Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
  • School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
  • Beijing Academy of Artificial Intelligence (BAAI)(北京智源人工智能研究院)
  • Tsinghua University(清华大学)
  • Amazon(亚马逊)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑