arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

保留未来,放弃展开:面向世界动作模型的RIFT

Keep the Future, Drop the Rollout: RIFT for World Action Models

Chushan Zhang, Jinguang Tong, Xuesong Li, Yikai Wang, Hongdong Li

arXiv 2608.11521首次发表:更新:

AI 中文总结

本文针对世界动作模型部署延迟问题,提出RIFT方法,通过学习的预期标记一次构建未来K/V缓存,在LIBERO、RoboTwin 2.0任务上实现高成功率且大幅降低动作块延迟。

AI 中文摘要

世界动作模型(WAMs)根据预测的未来来条件机器人动作,但迭代视频展开会增加部署延迟。本文探究动作生成是否需要不断变化的展开轨迹,还是仅需其未来表示。在全部40项LIBERO任务上针对4种WAMs开展的配对闭环干预实验显示,掩码或重新分配未来缓存值会改变执行效果并降低成功率,这表明模型对未来值及其分配位置具有敏感性。不过,对于Joint和Cosmos-2模型,重放一个固定的最终清理键/值(K/V)缓存几乎能保留未修改的执行效果,末端执行器平均位移误差为1.7至1.9厘米,成功率达97.9%至98.2%。这将缓存消耗与生产分离开来:这些模型可复用固定缓存,但仍需迭代展开来构建缓存。因此,本文提出RIFT(Rollout-free Imagination via Future Tokens,即通过未来标记实现无展开想象),该方法利用学习到的预期标记在一次骨干网络前向传播中构建完整的未来K/V缓存,同时保留原始的未来读取接口。在LIBERO任务上,RIFT的成功率达98.8%,与基于展开的Joint、IDM和LingBot-VA(成功率为98.4%至98.6%)接近,同时将动作块延迟降低了68.2%至89.1%。在RoboTwin 2.0上,RIFT在干净场景和随机场景中分别达到92.9%和92.6%的成功率,是评估方法中观测到的最高值。这些结果支持在部署时无需迭代视频生成即可实现无展开的未来条件设置。

英文摘要

World action models (WAMs) condition robot actions on predicted futures, but iterative video rollout increases deployment latency. We ask whether action generation requires the evolving rollout trajectory or only its future representation. Across four WAMs on 40 simulated robotic manipulation tasks, paired closed-loop interventions show that blocking access to the future cache or reassigning its values changes execution and reduces success. Yet in the evaluated co-denoising settings, reusing one fixed final-clean key/value (K/V) cache throughout action denoising nearly preserves unmodified execution, with $1.7$--$1.9$ cm end-effector average displacement error. Obtaining this cache still requires iterative video generation. We therefore propose RIFT (Rollout-free Imagination via Future Tokens), which uses learned anticipation tokens to construct a complete future K/V cache in one backbone pass. On LIBERO, RIFT achieves $98.8\%$ overall success, outperforming all evaluated rollout-based methods while yielding a $3.1$--$9.2\times$ inference speedup. Without further training, it achieves $81.1\%$ overall success on the out-of-distribution LIBERO-Plus benchmark, a $+9.7$ percentage-point improvement over the strongest evaluated baseline. On RoboTwin, it achieves $92.9\%$ and $92.6\%$ success on clean and randomized scenes, respectively, the highest among the evaluated methods. On real-world manipulation tasks, RIFT achieves $45.3\%$ average success, a $+6.0$ percentage-point improvement over Fast-WAM-Joint. These results support rollout-free future conditioning without iterative video generation at deployment.

CommentsAdded real-world experiments and updated the project URL

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑