发表机构
Fudan University; TARS Robotics; ShanghaiTech University; Institute of Automation, Chinese Academy of Sciences; Shanghai Jiao Tong University(复旦大学; TARS Robotics; 上海科技大学; 中国科学院自动化研究所; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出ΔWAM方法,通过动作切向量场将世界监督聚焦于动作相关动态,蒸馏进世界动作模型,在多个基准上提升鲁棒性并实现高效推理。
AI 中文摘要
世界动作模型(WAM)通过用密集的未来预测增强稀疏的动作监督来改进机器人策略。然而,大部分可预测的未来由外观和场景持续性主导,而非依赖于动作的动态。我们观察到,包括光流、运动中心表示和潜在动作在内的几种近期WAM设计,可以从一个共同视角来理解:当世界监督包含更高比例的动作相关变化时,其效率更高。基于这一见解,我们引入了动作切向量场,通过动作如何引起未来动态变化的局部泰勒展开来重新表述世界监督。我们在残差变分自编码器(Residual-VAE)空间中表示未来动态,其中未来潜变量可从当前潜变量及其残差中恢复,并使用一个强动作条件世界模型(ACWM)来探测动作变化与残差世界变化之间的局部对应关系。这种局部一阶结构被蒸馏到WAM中,以引导其去噪监督朝向与动作更紧密耦合的动态,而非仅可从外观预测的动态。在LIBERO-Plus、RoboTwin和RoboTwin2.0-Plus上,我们的方法持续提高了对光照、背景、相机、布局及其他环境扰动的鲁棒性。尽管未使用大规模具身预训练,它在多种分布偏移下比预训练策略实现了更强的鲁棒性。我们进一步将多步VideoDiT去噪蒸馏为单步,以实现高效推理。我们的结果表明,有效的WAM监督应保持信息丰富,同时将其预测能力集中在动作改变未来的方向上。
英文摘要
World Action Models (WAM) improve robot policies by augmenting sparse action supervision with dense future prediction. However, much of the predictable future is dominated by appearance and scene persistence rather than action-dependent dynamics. We observe that several recent WAM designs, including optical flow, motion-centric representations, and latent actions, can be understood from a common perspective in which world supervision becomes more efficient as it contains a higher proportion of action-relevant variation. Based on this insight, we introduce Action Tangent Fields, which reformulate world supervision through a local Taylor expansion of how actions induce changes in future dynamics. We represent future dynamics in Residual-VAE space, where the future latent remains recoverable from the current latent and its residual, and use a strong action-conditioned world model (ACWM) to probe the local correspondence between action variations and residual-world variations. This local first-order structure is distilled into the WAM to guide its denoising supervision toward dynamics that are more tightly coupled to action, rather than merely predictable from appearance. Across LIBERO-Plus, RoboTwin, and RoboTwin2.0-Plus, our method consistently improves robustness to lighting, background, camera, layout, and other environmental perturbations. Despite using no large-scale embodied pretraining, it achieves stronger robustness under several distribution shifts than pretrained policies. We further distill multi-step VideoDiT denoising into a single step for efficient inference. Our results suggest that effective WAM supervision should remain information-rich while concentrating its predictive capacity on the directions along which actions change the future.
Comments9 pages, 4 figures