发表机构
University of Maryland, College Park; Oxford Robotics Institute, University of Oxford(马里兰大学学院公园分校; 牛津大学牛津机器人研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出阶梯策略,一种流式推理与训练框架,将流匹配VLA转为JEPA式世界动作模型,通过交错去噪阶段分块执行动作,实现长时程高效推理,显著提升吞吐量并缩短首次动作时间。
AI 中文摘要
世界动作模型(WAMs)通过基于预测的未来观测来调节动作生成,从而改进机器人操作,但未来预测在已经昂贵的迭代动作生成之上增加了额外的推理开销。动作分块可以将此成本分摊到多个动作上,但在长执行时间范围内性能会下降,因为后续动作仍基于过时的观测进行条件化。我们引入了阶梯策略(STAIRCASE POLICY),一种流式推理和训练框架,它将流匹配VLA转变为JEPA风格的世界动作模型,并在交错的去噪阶段将大动作块划分为子块。近期动作一旦可用即被执行,而后续动作则继续被细化。在每个子块边界,未来潜变量根据最新观测重新预测,并用于更新所有未执行的动作,从而无需重复完整策略推理即可实现长时间范围执行。由此产生的未来预测误差可进一步用作自适应分块的信号。S-WAM在LIBERO上达到97.7%的准确率,在LIBERO-Plus上达到87.9%,并在多个策略主干和真实机器人任务上提升了性能。它每秒执行292.7个动作,是传统执行方式在相当精度下吞吐量的3.62倍,同时将首次动作时间从123.6毫秒缩短至73.3毫秒。通过额外的推理优化,吞吐量进一步增加至每秒642.9个动作。
英文摘要
World-Action Models (WAMs) improve robotic manipulation by conditioning action generation on predicted future observations, but future prediction adds further inference overhead to already expensive iterative action generation. Action chunking can amortize this cost over multiple actions, yet performance degrades over long execution horizons because later actions remain conditioned on stale observations. We introduce STAIRCASE POLICY, a streaming inference and training framework that turns a flow-matching VLA into a JEPA-style WAM and partitions a large action chunk into sub-chunks at staggered denoising stages. Near-term actions are executed as soon as they become available, while later actions continue to be refined. At each sub-chunk boundary, the future latent is re-predicted from the latest observation and used to update all unexecuted actions, enabling long-horizon execution without repeated full policy inference. The resulting future-prediction error can further serve as a signal for adaptive chunking. S-WAM achieves 97.7% on LIBERO and 87.9% on LIBERO-Plus, and improves performance across multiple policy backbones and real-robot tasks. It reaches 292.7 executed actions per second, $3.62\times$ the throughput of conventional execution at comparable accuracy, while reducing time-to-first-action from 123.6 to 73.3 ms. With additional inference optimizations, throughput further increases to 642.9 actions per second.