发表机构
Aether AI; University of California, San Diego(Aether AI; 加州大学圣地亚哥分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对动作条件世界模型学习需大量标记数据的问题,提出基于潜行动模型的因果去偏框架CD-LAM,通过三个微调目标改进模型,提升了潜行动可控性等多方面性能,减少了机器人动作适应更新次数。
AI 中文摘要
动作条件世界模型(ACWMs)旨在根据具身动作模拟未来观测,为机器人规划、策略评估和数据增强提供基础。但学习可控ACWMs需要大规模动作标记数据,成本高昂。潜行动模型(LAMs)通过从未标记视频中推断潜行动缓解瓶颈,但现有LAMs通常仅用重建目标训练,会将动作相关动态与背景等无关视觉因素纠缠。本文识别出这种与动作无关的偏差是可控ACWMs的关键障碍,引入评估指标。提出CD-LAM,一种基于LAM的ACWMs的因果去偏框架,引入三个微调目标。实验表明CD-LAM显著提高了潜行动可控性、下游机器人动作跟随、视觉保真度和适应效率。
英文摘要
Action-conditioned world models (ACWMs) aim to simulate future observations conditioned on embodied actions. However, learning controllable ACWMs requires large-scale action-labeled data, which remains costly to collect. Latent action models (LAMs) mitigate this bottleneck by inferring latent actions from videos without executable action labels, but existing LAMs are typically trained with reconstruction-only objectives and therefore entangle action-relevant dynamics with action-irrelevant visual confounders. With such objectives, LAM favors visual context over action dynamics, therefore downstream ACWMs exhibit residual motion under zero actions and fail to reproduce supplied dynamics. We propose CD-LAM, a causally debiased framework for LAM-based ACWMs. CD-LAM introduces three debiasing objectives: embodiment-centric reconstruction, action-centric contrastive learning, and latent space calibration, which together encourage embodiment-focused, action-aware, and well-calibrated representations. CD-LAM reduces action-following error by up to 42% and 35% when conditioned on latent and robot actions, while improving fidelity, and matching DreamDojo reference with over 12$\times$ fewer adaptation updates. CD-LAM also adapts LTX-2.3-22B into an ACWM with competitive quality using only 3.2% of DreamDojo-14B's sample exposures. On X-VLA, pretraining with CD-LAM instead of baseline LAM consistently improves success rates by up to 5 percent point across LIBERO, LIBERO-Plus, and RoboTwin C2R.