发表机构
Institute of Automation, Chinese Academy of Sciences; Galbot; Peking University; Shanghai Jiao Tong University(中国科学院自动化研究所; Galbot; 北京大学; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DualWAM通过解耦全局规划与局部细化,实现异步双系统世界动作模型,在保持长视野规划的同时提升闭环响应性,零样本操作成功率平均提升4.5个百分点,并实现16.6倍加速。
AI 中文摘要
世界动作模型(WAMs)联合生成机器人动作并预测未来世界状态,将视频预训练的先验知识迁移到机器人控制中。然而,未来视觉预测的计算成本高昂,因此现有WAMs通常依赖长动作块来分摊跨控制步骤的推理成本,但这以牺牲闭环响应性为代价。我们提出\method,一种双系统WAM,通过解耦全局规划与局部细化,在保持更广视野的世界动作生成的同时,实现高频闭环动作更新。\systwo周期性地在更广的世界动作块上执行高噪声双向去噪,以建立全局计划,而仅使用腕部相机的\sysone从中间去噪状态中提取时间对齐的短窗口,并利用最新的腕部观测完成低噪声细化,这些观测提供了交互过程中关于局部几何、运动和接触的动作对齐线索。两个系统沿共享去噪轨迹异步运行:每个全局计划被多次局部更新复用,而\sysone反复整合新的交互反馈。在Franka和Galbot上的零样本操作任务中,\method相比最强评估基线平均提高了4.5个百分点的成功率,同时实现了16.6倍的临界路径加速。进一步研究表明,角色匹配的自我中心和UMI数据将成功率提高了14个百分点,且解耦设计自然支持边缘-云部署,通信开销显著低于基线。
英文摘要
World Action Models (WAMs) jointly generate robot actions and predict future world states, transferring priors from video pretraining to robot control. However, future visual prediction is computationally expensive, so existing WAMs often rely on long action chunks to amortize inference cost across control steps, at the cost of closed-loop responsiveness. We present \method, a dual-system WAM that preserves broader-horizon world-action generation while enabling high-frequency closed-loop action updates by decoupling global planning and local refinement. \systwo periodically performs high-noise bidirectional denoising over a broader world-action chunk to establish a global plan, while wrist-only \sysone extracts a temporally aligned short window from the intermediate denoising state and completes low-noise refinement using the latest wrist observations, which provide action-aligned cues about local geometry, motion, and contact during interaction. The two systems operate asynchronously along a shared denoising trajectory: each global plan is reused across multiple local updates, while \sysone repeatedly incorporates fresh interaction feedback. Across zero-shot manipulation tasks on Franka and Galbot, \method improves success over the strongest evaluated baseline by 4.5 percentage points on average, while achieving a 16.6$\times$ critical-path speedup. Further studies show that role-matched egocentric and UMI data improve success by 14 percentage points, and that the decoupled design naturally supports edge--cloud deployment with substantially lower communication overhead than the baseline.
CommentsProject page: https://steveouo.github.io/DualWAM-Web/