arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

无视觉前瞻:面向世界动作模型的潜在未来

Foresight Without Seeing: Latent Futures for World Action Models

Jiakai Huang, Zhongbo Wu, Siyu Xu, Zheng Zhang, Zihan Wang, Shan You, Chang Xu, Tao Huang

arXiv 2608.11605首次发表:更新:

发表机构

Shanghai Jiao Tong University; ACE Robotics; Nanyang Technological University(上海交通大学; ACE机器人公司; 南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出ForeWAM,一种动力学条件下的直接策略WAM,通过Future-KV和动力学寄存器在无需显式生成未来视频的情况下,使直接策略WAMs兼具高效动作预测与预测动力学暴露能力,在LIBERO和LIBERO-Plus上取得高成功率。

AI 中文摘要

世界动作模型(World Action Models, WAMs)将未来视觉预测与机器人动作生成相结合,使策略能够对交互过程中物理世界的演化进行建模。现有WAMs在将预测动力学暴露给动作通路的方式上存在差异:显式未来WAMs可直接获取预测的场景演化信息,但迭代视频去噪会产生大量推理成本;相比之下,直接策略WAMs能从当前观测中高效预测动作,但缺乏在推理时将预测动力学暴露给Action DiT的接口。为弥合这一差距,本文提出ForeWAM,一种动力学条件下的直接策略WAM,无需解码未来视频即可为动作生成提供预测上下文。其核心是Future-KV,该模块在当前视觉隐态和随机未来槽上执行一次Video DiT预填充,并在整个动作去噪过程中复用生成的分层键值状态。我们还引入由冻结隐态动作教师监督的动力学寄存器,促使隐式未来状态捕获交互诱导的转变,如物体运动、接触变化和任务进展。真值未来观测和教师仅在训练期间使用,部署时无需这些,也不生成未来视频。在未进行具身机器人数据预训练的情况下,ForeWAM的标准变体和加速变体在LIBERO上分别达到96.7%和96.9%的平均成功率,标准变体在LIBERO-Plus上还达到61.6%的成功率。这些结果表明,直接策略WAMs可在保留高效动作预测的同时,无需显式生成未来观测即可将预测动力学暴露给动作通路。

英文摘要

World Action Models (WAMs) connect visual prediction with robot control, but supplying predictive context often requires expensive future-video generation. Direct policies avoid this cost but lack an explicit interface for accessing future-indexed predictive information. We introduce ForeWAM, a World Action Model that separates forecasting from rendering to expose and shape latent predictive context for efficient control. Its core mechanism, Future-KV, performs a single Video DiT prefill over the current visual latent and noise-initialized future slots, then reuses the resulting key-value states throughout action denoising. To make this context relevant to control, we introduce dynamics registers supervised by latent actions from a frozen teacher during training, encouraging representations of interaction-induced transitions. This reusable context supports a lightweight, single-layer action decoder. We evaluate ForeWAM on LIBERO, LIBERO-Plus, RoboCasa, and real-world manipulation tasks. Without additional policy-level embodied pretraining, ForeWAM improves RoboCasa success by 9.7 percentage points over Fast-WAM at the same budget of 50 demonstrations per task, reaching 59.2%. With a single-layer decoder, it achieves 77.6% success on LIBERO-Plus and reduces policy-query latency to 88.7 ms on an NVIDIA A800, delivering a 6.27-fold speedup over Fast-WAM. These results show that latent predictive computation provides useful foresight for robust, efficient control without explicit future-video generation.

Comments17 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑