发表机构
Johns Hopkins University; Tsinghua University; University of Chinese Academy of Sciences; Yinwang Intelligent Technology Co., Ltd.(约翰斯·霍普金斯大学; 清华大学; 中国科学院大学; 银望智能科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对异步世界-动作模型双时钟问题,提出ReSync方法,通过承诺-证据差距确定最优计算区间,在保持动作状态下推进世界流,无需参数调整,在RoboCasa上成功率提升4.48点。
AI 中文摘要
联合生成未来视频和动作已成为世界-动作模型的标准方法,最强系统在两个流上采用不同调度进行去噪:动作在少数步骤内解码以保持控制快速,而视频流运行更长时间以保持预测未来的清晰度。这种设计是刻意的,但使两个流处于不同时钟上,一个动作可能在支撑它的未来仍大体未解决时变得可执行。我们将此形式化为异步推理的双时钟视图,并引入承诺-证据差距,这是一个直接从模型自身采样调度读取而非通过搜索测量的量。该差距具有预测性:随着差距扩大,候选效用变得更难识别,额外候选采样收益减少,而推进世界流收益增加,两者交叉。因此,花费更多世界计算并非简单更好。有用区间两端封闭,两端都可在任何展开前从调度中读出。ReSync将计算置于该区间内:保持动作状态,仅在支持窗口内推进世界,然后恢复原生去噪。不改变参数,不比较候选。在冻结的配对RoboCasa面板上,这使成功率提高4.48个百分点,而等待但不推进世界的等计算对照组没有变化,同一规则无需重新调整即可迁移到第二个基准和第二个骨干网络。
英文摘要
Jointly generating future video and actions has become a standard recipe for world-action models, and the strongest systems denoise the two streams on separate schedules: actions are decoded in few steps so control stays fast, while the video stream runs longer to keep the predicted future sharp. The design is deliberate, but it leaves the two streams on different clocks, and an action can become executable while the future that should justify it is still largely unresolved. We formalize this as a two-clock view of asynchronous inference and introduce the commitment-evidence gap, a quantity read directly from a model's own sampling schedule rather than measured by search. The gap is predictive: as it widens, candidate utility becomes harder to identify and extra candidate sampling buys less, while advancing the world stream buys more, and the two cross. Spending more world computation is therefore not simply better. The useful interval is closed at both ends, and both ends can be read off the schedule before any rollout. ReSync places the computation inside it: hold the action state, advance only the world within the supported window, then resume native denoising. No parameters change and no candidates are compared. On a frozen paired RoboCasa panel this improves success by 4.48 points, while an equal-compute control that waits without advancing the world does not move, and the same rule transfers to a second benchmark and a second backbone without retuning.