发表机构
Institute for AI Industry Research (AIR), Tsinghua University; Shanghai Jiao Tong University; The Hong Kong University of Science and Technology (Guangzhou); AIR Wuxi Innovation Center, Tsinghua University; The University of Adelaide; Wuhan University; Southeast University; Beijing Jiaotong University; Fudan University; Li Auto; School of Information, Renmin University of China(清华大学人工智能产业研究院(AIR); 上海交通大学; 香港科技大学(广州); 清华大学AIR无锡创新中心; 阿德莱德大学; 武汉大学; 东南大学; 北京交通大学; 复旦大学; 理想汽车; 中国人民大学信息学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MobileWAM是一种混合Transformer架构,通过前瞻链(CoF)解决移动操作的异质动态问题,在ManiSkill-HAB上超越SOTA策略,可微调至真实移动操作机器人且泛化性强。
AI 中文摘要
基于视频生成主干构建的世界动作模型(WAMs)是机器人学习领域的新兴方法,但目前仅局限于桌面操作任务。移动操作需要在场景级动态环境中同时完成移动和全身操作,目前仍依赖于动态盲的视觉编码器与手工设计的协调机制。我们通过MobileWAM弥合了这一差距,它是一种混合Transformer架构,通过分层联合注意力将预训练的视频扩散Transformer与轻量级动作专家融合,将互联网级运动先验转化为全身控制。为协调移动与操作的异质动态,动作专家的每个前馈层成为共享、移动和操作三个专家的混合体,由动作令牌中的运动意图进行软路由。为了强化监督,我们进一步提出了前瞻链(CoF):中间表示按顺序预测一系列未来潜在块,每一步都以前一步为条件。CoF与我们解耦的视频-动作去噪方案自然适配。在部署时,WAM仅作为当前帧编码器;前瞻仅通过梯度发挥作用,因此推理时会丢弃前瞻链和视频生成,仅保留策略级成本。MobileWAM在ManiSkill-HAB上超越了最先进的移动操作策略,并在真实ARX Lift2移动操作机器人上针对不同任务进行微调,展现出强大的泛化能力。代码将很快发布。
英文摘要
World action models (WAMs) built on video generation backbones are a rising recipe for robot learning, yet remain confined to tabletop manipulation. Mobile manipulation demands simultaneous locomotion and whole-body manipulation amid scene-scale dynamics, yet is still dominated by dynamics-blind visual encoders with hand-crafted coordination. We bridge this gap with MobileWAM, a mixture-of-transformers architecture that fuses a pretrained video diffusion transformer with a lightweight action expert through layerwise joint attention, translating internet-scale motion priors into whole-body control. To reconcile the heterogeneous dynamics of moving and manipulating, each feed-forward layer of the action expert becomes a three-expert mixture of shared, locomotion, and manipulation experts, softly routed by the motion intent in the action tokens. To densify supervision, we further propose Chain-of-Foresight (CoF): intermediate representations sequentially predict a chain of future latent chunks, each step conditioned on its predecessor. CoF pairs naturally with our decoupled video--action denoising scheme. At deployment, the WAM serves as a pure current-frame encoder; foresight acts only through gradients, so at inference the foresight chain and video generation are discarded, leaving only policy-level cost. MobileWAM surpasses state-of-the-art mobile manipulation policies on ManiSkill-HAB and fine-tunes to a real ARX Lift2 mobile manipulator across diverse tasks with strong generalization. Code will be released soon.