AI 中文总结
本文针对世界动作模型(WAMs)的固定执行 horizon 缺陷,提出自适应执行方案TempoWAM,经实验可改善WAM执行的效率-成功率权衡,在真实机器人任务中表现优异。
AI 中文摘要
世界动作模型(World Action Models, WAMs)可联合预测未来动作与环境演化。每次推理时,WAM会生成一段动作序列,机器人执行固定前缀后再重新规划。本文指出这种固定执行 horizon 与执行动态不匹配:不同任务阶段的序列可靠性存在差异,何时重新规划取决于累积执行结果,而非步数。我们提出TempoWAM(Timing Execution by Monitoring Progress Online,即通过在线监控进度确定执行时机),一种适用于WAM的轻量即插即用执行方案。循环进度监控器会从当前观测、任务指令、剩余动作及执行历史中估计任务进度;自适应执行协议则评估该序列是否在推进任务,以决定是否需要重新规划。为弥合训练-部署差距,该协议会通过任务相关的校准因子进行在线调整。在LIBERO、RoboTwin及真实任务上的实验表明,TempoWAM可持续改善WAM执行的效率-成功率权衡。在真实机器人上,其在简单任务中减少26.9%的WAM推理次数且保持成功率,在困难任务中则提升13.3个百分点的成功率。
英文摘要
World Action Models (WAMs) jointly predict future actions and the evolution of the environment. At each inference, a WAM generates a chunk of actions and the robot executes a fixed prefix before replanning. We argue that this fixed execution horizon is poorly matched to execution dynamics: the chunk reliability varies across task stages, so when to replan depends on the result of accumulated execution, not on the step counts. We propose TempoWAM (Timing Execution by Monitoring Progress Online), a lightweight plug-and-play execution scheme for WAMs. A Recurrent Progress Monitor first estimates task progress from the current observation, task instruction, remaining actions, and execution history; and an Adaptive Execution Protocol then evaluates whether the chunk is advancing the task to decide if replanning is needed. To bridge the training-deployment gap, the protocol is calibrated by a task-dependent calibration factor with online adaptation. Experiments on LIBERO, RoboTwin, and real-world tasks show that TempoWAM consistently improves the efficiency-success trade-off of WAM execution. On real robots, it reduces WAM inferences by 26.9% on easy tasks while maintaining success, and improves success by 13.3 points on difficult tasks.