AI 中文总结
本研究提出模型无关的4D-WAM,通过运动对齐与目标对齐两个互补目标,将3D轨迹场的时空知识注入世界动作模型,在多基准实验中显著提升了模型的空间理解等多项性能。
AI 中文摘要
基于世界模型的最新进展,世界动作模型(World Action Models, WAMs)可联合建模视频预测与动作生成。然而,这类模型通常在2D像素空间中表示视频,与机器人动作执行所处的3D空间存在表示差距。近期的3D方法虽引入了3D信息,但未能充分利用3D结构的动态特性。本研究提出4D-WAM,一种模型无关的训练策略,通过表示对齐将来自3D轨迹场的时空知识注入WAMs。为此,我们引入两个互补目标:1)运动对齐,对齐相邻帧间的时间特征变化,鼓励模型在训练过程中建立局部4D感知;2)目标对齐,通过最小化源帧与目标帧间类注意力相似度分布的差距,引导模型从源帧推断最终目标。这两个目标共同提供局部运动监督与长时程目标引导,使WAMs能够学习轨迹级的时空表示。在不同基础模型上开展的大量分布内与分布外实验表明,该模型在空间理解、执行精度、鲁棒性、泛化性及通用性方面均有提升。
英文摘要
Building on recent advances in world models, World Action Models (WAMs) jointly model video prediction and action generation. However, they typically represent videos in 2D pixel space, creating a representation gap with 3D space in which robotic actions are executed. Recent 3D approaches introduce 3D information, but fail to fully exploit the dynamics of 3D structures. In this work, we propose 4D-WAM, a model-agnostic training strategy that injects spatiotemporal knowledge from 3D trajectory fields into WAMs through representation alignment. To this end, we introduce two complementary objectives: 1) motion alignment, which aligns temporal feature variations across adjacent frames and encourages the model to build local 4D awareness during training, and 2) destination alignment, which guides the model to infer the final destination from the source frame by minimizing the gap between their attention-like similarity distributions. Together, these objectives provide both local motion supervision and long-horizon goal guidance, enabling WAMs to learn trajectory-level spatiotemporal representations. Extensive in-distribution and out-of-distribution experiments across different base models demonstrate the model's improvements in spatial understanding, execution precision, robustness, generalization, and versatility.