发表机构
Huawei Noah’s Ark Lab; Huawei Celia Team; Labs(华为诺亚方舟实验室; 华为Celia团队; 2012实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有WAMs动作模块深度绑定视频骨干网络导致延迟高的问题,提出以视频为中心的DoT架构,构建Faster-WAM,实现低延迟、高性能与强泛化,较Fast-WAM提速3.2倍。
AI 中文摘要
世界动作模型(World Action Models, WAMs)将机器人动作预测与视频世界模型相结合。现有采用共享骨干网络和混合专家Transformer设计的WAMs通常将动作模块的深度与视频骨干网络的深度绑定,导致大量计算开销和高推理延迟。为解决这一局限,我们提出Transformer对接(Dock of Transformer, DoT),这是一种以视频为中心的设计原则,将预训练视频Transformer作为表示中心,通过对接接口连接轻量级输出头。该设计支持灵活的输出头设计,同时可直接访问骨干网络所有层的表示。随后我们提出Faster-WAM,即DoT在WAMs中的实例,其将单层动作头对接至30层视频骨干网络;对接接口融合所有视频层的键与值,并应用旋转位置嵌入(RoPE)重新对齐。无需额外具身预训练,Faster-WAM在LIBERO和RoboTwin 2.0数据集上实现了具有竞争力的性能,同时在LIBERO-Plus上展现出强大的分布外泛化能力。Faster-WAM在我们的受控对比中还实现了最低的端到端延迟,每次推理仅需66.5毫秒,较Fast-WAM提速3.2倍。总体而言,这些结果表明以视频为中心的DoT架构支持灵活的任务特定头设计,同时兼具低推理延迟、出色的动作预测性能和稳健的泛化能力。
英文摘要
World Action Models (WAMs) build on pretrained video models, whose representations are grounded in physical dynamics and provide a natural basis for action prediction. Despite this natural foundation, many WAMs still rely on deep, parameter-heavy action-prediction modules that incur high inference latency and may overfit to limited robot demonstrations, restricting their real-world applicability. In this paper, we advocate a world-model-centric principle that concentrates capacity and computation in the video world model, while a lightweight action expert translates the backbone's representations into executable robot actions. We realize this principle through three key choices: Dock of Transformers (DoT) with Lite KV-Fusion to give the shallow, lightweight action expert access to representations from all video layers; world-model-only conditioning of the action expert; and retracted 1D-RoPE for positional alignment between video keys and action queries. We test this principle using Faster-WAM, a world-model-centric WAM with only a single-layer action expert. Despite this restriction on action-specific computation, Faster-WAM achieves competitive control performance on LIBERO and RoboTwin~2.0 without additional embodied pretraining. It provides approximately $3.7\times$ and $1.3\times$ inference speedups over Fast-WAM and $π_{0.5}$, respectively. Consistent with its world-model-centric design, Faster-WAM demonstrates stronger generalizability under distribution shifts: the same LIBERO-trained policy achieves $78.3\%$ success on LIBERO-Plus, exceeding Fast-WAM and LingBot-VA by $26.8$ and $8.8$ percentage points, respectively. Finally, real-robot experiments demonstrate success rates comparable to Fast-WAM, with substantially lower inference latency and shorter task-completion times.