arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MobileWAM:用前瞻链将世界动作模型(WAM)与移动操作任务对接

MobileWAM: Bridging World Action Models to Mobile Manipulation with Chain-of-Foresight

Zehua Fan, Junjie He, Wenxuan Song, Xi Wang, Wenqi Lyu, Linge Zhao, Fuhao Li, Zihan You, Yifei Yang, Kaiming Xu, Qi Jiang, Yue Jiang, Haoang Li, Cheng Chi, Feng Gao, Bailin Li, Yan Wang

arXiv 2608.04657首次发表:更新:

发表机构

Institute for AI Industry Research (AIR), Tsinghua University; Shanghai Jiao Tong University; The Hong Kong University of Science and Technology (Guangzhou); AIR Wuxi Innovation Center, Tsinghua University; The University of Adelaide; Wuhan University; Southeast University; Beijing Jiaotong University; Fudan University; Li Auto; School of Information, Renmin University of China(清华大学人工智能产业研究院(AIR); 上海交通大学; 香港科技大学(广州); 清华大学AIR无锡创新中心; 阿德莱德大学; 武汉大学; 东南大学; 北京交通大学; 复旦大学; 理想汽车; 中国人民大学信息学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MobileWAM是一种混合Transformer架构,通过前瞻链(CoF)解决移动操作的异质动态问题,在ManiSkill-HAB上超越SOTA策略,可微调至真实移动操作机器人且泛化性强。

AI 中文摘要

基于视频生成主干构建的世界动作模型(WAMs)是机器人学习领域的新兴方法,但目前仅局限于桌面操作任务。移动操作需要在场景级动态环境中同时完成移动和全身操作,目前仍依赖于动态盲的视觉编码器与手工设计的协调机制。我们通过MobileWAM弥合了这一差距,它是一种混合Transformer架构,通过分层联合注意力将预训练的视频扩散Transformer与轻量级动作专家融合,将互联网级运动先验转化为全身控制。为协调移动与操作的异质动态,动作专家的每个前馈层成为共享、移动和操作三个专家的混合体,由动作令牌中的运动意图进行软路由。为了强化监督,我们进一步提出了前瞻链(CoF):中间表示按顺序预测一系列未来潜在块,每一步都以前一步为条件。CoF与我们解耦的视频-动作去噪方案自然适配。在部署时,WAM仅作为当前帧编码器;前瞻仅通过梯度发挥作用,因此推理时会丢弃前瞻链和视频生成,仅保留策略级成本。MobileWAM在ManiSkill-HAB上超越了最先进的移动操作策略,并在真实ARX Lift2移动操作机器人上针对不同任务进行微调,展现出强大的泛化能力。代码将很快发布。

英文摘要

World action models (WAMs) built on video generation backbones are a rising recipe for robot learning, yet remain confined to tabletop manipulation. Mobile manipulation demands simultaneous locomotion and whole-body manipulation amid scene-scale dynamics, yet is still dominated by dynamics-blind visual encoders with hand-crafted coordination. We bridge this gap with MobileWAM, a mixture-of-transformers architecture that fuses a pretrained video diffusion transformer with a lightweight action expert through layerwise joint attention, translating internet-scale motion priors into whole-body control. To reconcile the heterogeneous dynamics of moving and manipulating, each feed-forward layer of the action expert becomes a three-expert mixture of shared, locomotion, and manipulation experts, softly routed by the motion intent in the action tokens. To densify supervision, we further propose Chain-of-Foresight (CoF): intermediate representations sequentially predict a chain of future latent chunks, each step conditioned on its predecessor. CoF pairs naturally with our decoupled video--action denoising scheme. At deployment, the WAM serves as a pure current-frame encoder; foresight acts only through gradients, so at inference the foresight chain and video generation are discarded, leaving only policy-level cost. MobileWAM surpasses state-of-the-art mobile manipulation policies on ManiSkill-HAB and fine-tunes to a real ARX Lift2 mobile manipulator across diverse tasks with strong generalization. Code will be released soon.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑