发表机构
NVIDIA; MIT; HKU; UCSD(英伟达; 麻省理工学院; 香港大学; 加州大学圣迭戈分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Long-WAM通过自回归预训练和流式系统设计,扩展因果世界-动作模型的上下文,在实时控制中显著提升成功率,并支持动态长时程操作。
AI 中文摘要
实时机器人控制需要足够的视觉历史来推断运动和任务进度,但处理这些历史可能会延迟动作执行。我们提出了Long-WAM,一个在实时控制约束下扩展因果世界-动作模型上下文的模型-系统框架。我们的核心发现是,获取历史与使用历史并不相同:当视频基础模型通过自回归(AR)方式进行预训练时,更长的历史带来的收益要大得多。我们首先在没有动作标签的情况下从机器人和第一人称视频中学习因果预测,然后在世界-动作适应过程中保留这种从历史到未来的结构。在RoboCasa GR-1上,将上下文从0.0秒增加到19.2秒,成功率从63.3%提升到78.7%,而双向预训练初始化则没有显示出净收益;机器人领域的AR预训练进一步提高了GR-1和LIBERO-Long上的峰值成功率。Long-WAM在LIBERO-Long、RoboTwin 2.0和DOMINO上也取得了比较方法中的最佳结果。流式观测编码、异步执行和硬件特定加速使其能够在RTX 5090、DGX Spark和Jetson AGX Thor上部署,而不会丢失未来预测;在RTX 5090上,每个动作块(包括未来视频潜在预测)耗时107.4毫秒。在Unitree G1和YAM上的实时部署支持动态和长时程操作,包括在动态叠杯子任务中达到95%的成功率,而Pi0.5和Fast-WAM在20次试验中均未成功。作为具有记忆能力的执行器,Long-WAM在复合任务中还能补充高层规划。
英文摘要
Real-time robot control demands enough visual history to infer motion and task progress, but processing that history can delay action. We present Long-WAM, a model-system framework for scaling the context of causal world-action models under real-time control constraints. Our central finding is that access to history is not the same as using it: longer histories pay off far more when the video foundation is pretrained autoregressively (AR). We first learn causal prediction from robot and egocentric videos without action labels, then preserve this history-to-future structure during world-action adaptation. On RoboCasa GR-1, increasing context from 0.0 to 19.2 seconds raises success from 63.3% to 78.7%, whereas a bidirectionally pretrained initialization shows no net gain; robot-domain AR pretraining further raises peak success on GR-1 and LIBERO-Long. Long-WAM also achieves the best results among compared methods on LIBERO-Long, RoboTwin 2.0, and DOMINO. Streaming observation encoding, asynchronous execution, and hardware-specific acceleration enable deployment on RTX 5090, DGX Spark, and Jetson AGX Thor without dropping future prediction; on RTX 5090, each action chunk, including future-video latent prediction, takes 107.4 ms. Real-time deployment on Unitree G1 and YAM supports dynamic and long-horizon manipulation, including 95% success on dynamic cup stacking, where Pi0.5 and Fast-WAM succeed in none of 20 trials. As a memory-informed executor, Long-WAM also complements higher-level planning in composite tasks.