发表机构
CollegeAI, THU; AIR, THU; CASIA; TARS Robotics; FDU; Southeast University; SJTU; NUS; BUAA(清华大学CollegeAI; 清华大学AIR; 中国科学院自动化研究所; TARS机器人公司; 复旦大学; 东南大学; 上海交通大学; 新加坡国立大学; 北京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出名为 LAWA 的 WAM 架构,以紧凑隐式动作作为未来意图实现高效未来想象,在 RoboCasa 等数据集上优于 Fast-WAM,延迟更低且性能更优,兼顾性能、泛化性与延迟。
AI 中文摘要
世界动作模型(World Action Models, WAMs)通过对观测的演化过程建模来提升机器人控制性能,但在测试时生成未来观测会带来显著的延迟。Fast-WAM 为提升效率去除了这一过程;然而,我们的匹配实现显示,Fast-WAM 的泛化能力低于具有未来感知的替代方案,尤其是在机器人演示数据稀缺和分布外场景下。为弥合这一差距,我们提出了 LAWA,一种 WAM 架构,它将紧凑的隐式动作作为未来意图的操作表示,无需生成未来观测即可实现高效的测试时未来想象。具体而言,一种通过无动作预训练增强的离散分词器生成以操纵为中心的码本目标。LAWA 在推理时省略未来视频分支,联合对锚定在这些目标上的连续隐式状态与可执行动作块进行去噪。在 RoboCasa 上,LAWA 在少样本和全数据设置中分别达到 65.6% 和 80.8% 的平均成功率,较匹配的 Fast-WAM 基线分别提升 9.6 和 4.5 个百分点。它还保持了匹配的 Joint-WAM 变体的性能水平,同时推理延迟降低 42.9%。LAWA 在 LIBERO-Plus 上也展现出有竞争力的零样本鲁棒性,并在真实世界任务上表现出更优性能。这些结果表明,未来想象无需被舍弃:通过紧凑的隐式动作保留未来想象,可在性能、泛化性和延迟之间实现有效权衡。代码和模型将被发布。
英文摘要
World action models (WAMs) improve robot control by modeling how observations evolve, but generating future observations at test time incurs substantial latency. Fast-WAM removes this process for efficiency; however, our matched implementations show lower generalization for Fast-WAM than for future-aware alternatives, especially with scarce robot demonstrations and in out-of-distribution scenarios. To bridge this gap, we introduce **LAWA**, a WAM architecture that uses compact latent actions as an operational representation of future intentions, enabling efficient test-time future imagination without generating future observations. Specifically, a discrete tokenizer enhanced by action-free pre-training produces manipulation-centric codebook targets. LAWA jointly denoises a continuous latent state anchored to these targets with executable action chunks while omitting the future-video branch at inference. On RoboCasa, LAWA achieves state-of-the-art average success rates of 65.6% and 80.8% in the few-shot and full data settings, improving over the matched Fast-WAM baseline by 9.6 and 4.5 points, respectively. It also preserves the performance level of the matched Joint-WAM variant while requiring 42.9% lower inference latency. LAWA also demonstrates competitive zero-shot robustness on LIBERO-Plus and superior performance on real-world tasks. These results show that future imagination need not be discarded: retaining it with compact latent actions yields an effective trade-off among performance, generalization, and latency. Code and models will be released.