发表机构
Virginia Tech; Drexel University; Northeastern University; Purdue University(弗吉尼亚理工大学; 卓克索大学; 东北大学; 普渡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出GlanceWAM,通过异步解耦视频DiT的想象与控制,打破世界-动作模型的速度-成功率困境,在RoboCasa、LIBERO基准测试中表现优异,推理速度达48ms/块。
AI 中文摘要
视频生成模型为机器人学习提供了丰富的物理先验,但现有的世界-动作模型(WAMs)面临一个根本的权衡:以控制频率同步生成视频会因延迟过高而无法使用,而放弃测试时的视觉想象则会降低任务成功率。我们表明,当视觉想象在关键路径外异步生成并直接在潜空间中使用时,可同时实现实时推理和更高的成功率。我们提出了GlanceWAM,它在单个视频DiT中将想象与控制解耦:异步提议者以较慢的时钟速率在后台超前想象未来数秒的单个前瞻帧,而动作头则以控制频率(48毫秒)仅在潜空间中解码动作块,不会造成阻塞。通过采用隔离视频表示的无干扰注意力掩码,以及适应异步前瞻老化的抗陈旧性地平线训练,GlanceWAM打破了速度-成功率困境。它仅通过演示进行训练,在包含24个任务的RoboCasa厨房基准测试中达到72.2%的成功率(超过同步Cosmos Policy的67.1%和无想象协同训练的64.4%),在LIBERO基准测试中达到99.0%的成功率,在NVIDIA A100 GPU上的执行速度为每个块48毫秒(比同步基准快24倍)。代码可在此https URL获取。
英文摘要
Video generative models provide rich physical priors for robot learning, yet existing world-action models (WAMs) face a fundamental trade-off: synchronous video generation at control rate is latency-prohibitive, while abandoning test-time visual imagination sacrifices task success. We show that visual imagination achieves both real-time inference and superior success rates when generated asynchronously off the critical path and consumed directly in latent space. We introduce GlanceWAM, which decouples imagination from control on a single shared video DiT backbone: an asynchronous proposer glances ahead on a slow clock to imagine a single lookahead frame seconds into the future in the background, while an action head decodes action chunks at control rate (48 ms) purely in latent space without blocking. Enabled by a non-interfering attention mask that isolates video representations and staleness-robust horizon training that accommodates asynchronous lookahead aging, GlanceWAM breaks the speed-success dilemma. Trained purely on demonstrations, it attains 72.2% on the 24-task RoboCasa kitchen benchmark (vs. 67.1% for synchronous Cosmos Policy) and 99.0% on LIBERO while cutting per-chunk control latency $24\times$ relative to synchronous world-action models (48 ms on one A100). In single-arm and bimanual real-robot manipulation, it achieves higher average success than $π_{0.5}$ without any robot-data pretraining. Code is available at https://github.com/linhanwang/GlanceWAM.
CommentsAdd real-robot experiments