arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Faster-WAM:世界动作模型需要深度动作模块吗?

Faster-WAM: Do World Action Models Need Deep Action Modules?

Liheng Ma, Rui Heng Yang, Amin Abyaneh, George Z. Xue, Behnam Rahmati, Mateo Clemente, Ziwen Hu, Anlin Chen, Tongtong Cao, Zhanguang Zhang, Yingxue Zhang

arXiv 2608.02365首次发表:更新:

发表机构

Huawei Noah’s Ark Lab; Huawei Celia Team; Labs(华为诺亚方舟实验室; 华为Celia团队; 2012实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有WAMs动作模块深度绑定视频骨干网络导致延迟高的问题,提出以视频为中心的DoT架构,构建Faster-WAM,实现低延迟、高性能与强泛化,较Fast-WAM提速3.2倍。

AI 中文摘要

世界动作模型(World Action Models, WAMs)将机器人动作预测与视频世界模型相结合。现有采用共享骨干网络和混合专家Transformer设计的WAMs通常将动作模块的深度与视频骨干网络的深度绑定,导致大量计算开销和高推理延迟。为解决这一局限,我们提出Transformer对接(Dock of Transformer, DoT),这是一种以视频为中心的设计原则,将预训练视频Transformer作为表示中心,通过对接接口连接轻量级输出头。该设计支持灵活的输出头设计,同时可直接访问骨干网络所有层的表示。随后我们提出Faster-WAM,即DoT在WAMs中的实例,其将单层动作头对接至30层视频骨干网络;对接接口融合所有视频层的键与值,并应用旋转位置嵌入(RoPE)重新对齐。无需额外具身预训练,Faster-WAM在LIBERO和RoboTwin 2.0数据集上实现了具有竞争力的性能,同时在LIBERO-Plus上展现出强大的分布外泛化能力。Faster-WAM在我们的受控对比中还实现了最低的端到端延迟,每次推理仅需66.5毫秒,较Fast-WAM提速3.2倍。总体而言,这些结果表明以视频为中心的DoT架构支持灵活的任务特定头设计,同时兼具低推理延迟、出色的动作预测性能和稳健的泛化能力。

英文摘要

World Action Models (WAMs) build on pretrained video models, whose representations are grounded in physical dynamics and provide a natural basis for action prediction. Despite this natural foundation, many WAMs still rely on deep, parameter-heavy action-prediction modules that incur high inference latency and may overfit to limited robot demonstrations, restricting their real-world applicability. In this paper, we advocate a world-model-centric principle that concentrates capacity and computation in the video world model, while a lightweight action expert translates the backbone's representations into executable robot actions. We realize this principle through three key choices: Dock of Transformers (DoT) with Lite KV-Fusion to give the shallow, lightweight action expert access to representations from all video layers; world-model-only conditioning of the action expert; and retracted 1D-RoPE for positional alignment between video keys and action queries. We test this principle using Faster-WAM, a world-model-centric WAM with only a single-layer action expert. Despite this restriction on action-specific computation, Faster-WAM achieves competitive control performance on LIBERO and RoboTwin~2.0 without additional embodied pretraining. It provides approximately $3.7\times$ and $1.3\times$ inference speedups over Fast-WAM and $π_{0.5}$, respectively. Consistent with its world-model-centric design, Faster-WAM demonstrates stronger generalizability under distribution shifts: the same LIBERO-trained policy achieves $78.3\%$ success on LIBERO-Plus, exceeding Fast-WAM and LingBot-VA by $26.8$ and $8.8$ percentage points, respectively. Finally, real-robot experiments demonstrate success rates comparable to Fast-WAM, with substantially lower inference latency and shorter task-completion times.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑