AI 中文总结
本研究针对现有潜在动作模型缺乏3D感知能力的问题,提出LAWM-3D模型,通过多视图动作 token 化、几何对齐约束和RGB-D联合重建目标,实现了性能SOTA的机器人世界模型。
AI 中文摘要
世界模型使智能体能够在无需与真实世界交互的情况下执行前向滚动和规划,然而,其在开放世界具身智能中的应用仍受限于动作标注的高成本以及不同平台间动作空间的异质性。近年来,潜在动作模型(LAM)通过以自监督方式直接从未标注的人类视频中学习动作表示,缓解了这一瓶颈。不过,现有大多数LAM依赖单视图输入,且主要在2D像素空间中运行,这引发了一个根本性问题:仅将多视图视频纳入LAM训练是否能赋予所学潜在动作3D感知能力?本研究表明答案是否定的,主要原因在于未来帧外观泄漏,以及相机间的外观差异和视角变化。为解决这些问题,我们提出了LAWM-3D,引入了三个紧密耦合的关键设计:(1)用于学习3D感知潜在动作的多视图不变统一动作 token 化方案;(2)将中间编码器特征锚定到预训练3D基础模型的几何对齐约束,从而明确提供跨视图几何对应关系;(3)非单射RGB-D联合重建目标,以防止从未来帧外观信息中进行捷径学习,迫使LAM将监督集中于具有几何意义的运动线索。重要的是,这些组件并非简单堆叠,而是通过统一动机紧密耦合。基于大规模人类视频预训练后进行机器人微调的两阶段范式,大量实验表明,所提出的3D感知潜在动作显著提升了世界模型性能,在生成质量、物理一致性和泛化能力方面达到了SOTA结果。
英文摘要
World models enable agents to perform forward rollout and planning without real-world interaction. However, their application in open-world embodied intelligence remains limited by the high cost of action annotations and the heterogeneity of action spaces across platforms. Recently, latent action models (LAMs) have alleviated this bottleneck by learning action representations directly from unlabeled human videos in a self-supervised manner. Nevertheless, most existing LAMs rely on single-view inputs and operate primarily in 2D pixel space, raising a fundamental question: can simply incorporating multi-view videos into LAM training endow the learned latent actions with 3D-aware perception? Our study shows that the answer is negative. The primary reasons lie in future-frame appearance leakage as well as inter-camera appearance discrepancies and viewpoint variations. To address these issues, we propose LAWM-3D, which introduces three tightly coupled key designs: (1) a multi-view invariant unified action tokenization scheme for learning 3D-aware latent actions; (2) a geometric alignment constraint that anchors intermediate encoder features to a pretrained 3D foundation model, thereby explicitly providing cross-view geometric correspondences; and (3) a non-injective RGB-D joint reconstruction objective that prevents shortcut learning from future-frame appearance information, forcing the LAM to focus supervision on motion cues with geometric significance. Importantly, these components are not simply stacked but are tightly coupled through a unified motivation. Built upon a two-stage paradigm of large-scale human video pretraining followed by robot fine-tuning, extensive experiments demonstrate that the proposed 3D-aware latent actions significantly improve world model performance, achieving SOTA results in generation quality, physical consistency, and generalization ability.