发表机构
University of Southern California; Brown University; Fudan University; Toyota Research Institute(南加州大学; 布朗大学; 复旦大学; 丰田研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Rolling-WAM通过滚动想象将联合去噪分布到连续重新规划周期,实现4.5倍稳态加速,并在LIBERO、RoboTwin和Unitree G1上保持竞争力。
AI 中文摘要
世界动作模型(WAMs)将动作生成与未来视觉预测相结合,用于机器人操作。然而,在每个重新规划周期内完成联合视频-动作去噪过程会产生大量延迟,从而延迟动作更新并限制闭环响应能力。我们提出了Rolling-WAM,一种将联合去噪分布到连续重新规划周期中的公式化方法。我们的方法维护一个具有交错噪声水平的视频-动作块滑动窗口。在每一步中,滚动噪声调度完全去噪即将执行的动作块,同时部分细化更远的未来块。随着窗口随着新的相机观察而推进,保留的未来块继续其去噪过程。这将在时间上分布计算成本,同时跨块边界携带不断演变的视觉-动作上下文。在LIBERO、RoboTwin和真实世界的Unitree G1人形机器人上的评估表明,Rolling-WAM实现了具有竞争力的操作性能。通过消除从头开始对整个预测范围进行去噪的需要,它相对于标准联合WAM提供了4.5倍的稳态重新规划加速。
英文摘要
World Action Models (WAMs) couple action generation with future visual prediction for robotic manipulation. However, completing the joint video-action denoising process at each replanning cycle incurs substantial latency, delaying action updates and limiting closed-loop responsiveness. We present Rolling-WAM, a formulation that distributes joint denoising across successive replanning cycles. Our method maintains a sliding window of video-action chunks at staggered noise levels. At each step, a rolling noise schedule fully denoises the imminent action chunk for execution, while partially refining farther-future chunks. As the window advances with new camera observations, the retained future chunks continue their denoising process. This distributes the computational cost over time while carrying an evolving visual-action context across chunk boundaries. Evaluations on LIBERO, RoboTwin, and a real-world Unitree G1 humanoid show that Rolling-WAM achieves competitive manipulation performance. By removing the need to denoise the entire prediction horizon from scratch, it delivers a 4.5x steady-state replanning speedup over standard joint WAMs.
Comments10 pages, 7 figures, 5 tables. Under review. Project page: https://rolling-wam.github.io/