发表机构
Institute of Science Tokyo; National Institute of Informatics; DENSO IT Lab(东京科学大学; 信息学研究所; 电装IT实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对世界动作模型(WAMs)因未来序列选择影响性能的问题,提出世界一致解码(WCD)框架,通过采样排序候选并利用想象-现实不匹配训练在线预测器,在RoboTwin 2.0上提升了任务成功率与鲁棒性。
AI 中文摘要
世界动作模型(World Action Models, WAMs)旨在通过随机生成视觉未来序列再解码动作来控制机器人,但实证观察表明,结果高度依赖所选的未来序列。我们提出世界一致解码(World-Coherent-Decoding, WCD),一种自验证测试时规划框架,将WAM的滚动输出视为可证伪的未来-动作假设。在每个决策步骤,WCD从冻结的WAM中采样多个候选,并利用内部生成信号对其排序:基于流的视频意外性(用于视觉合理性)和动作路径代价(用于动作生成稳定性)。执行后,实际观测会校验所选的想象结果,产生想象-现实不匹配,用于训练轻量级在线预测器以进行未来候选选择。因此,WCD无需更新主干模型,即可将延迟的自验证转化为执行前的可靠性估计。在RoboTwin 2.0上,WCD在有限随机场景监督下将Hard任务成功率从55.80%提升至60.90%,在Horizon-3任务上实现了+16.43的增益,并在真实Franka视觉偏移测试中展现出定性鲁棒性。这些结果凸显了一个简单原则:WAM的测试时扩展更多依赖于选择可靠的未来序列,而非采样更多未来序列。
英文摘要
World Action Models (WAMs) aim to control robots by stochastically generating visual futures and then decoding actions, but empirical observations indicate that the results can strongly depend on which future is selected. We propose World-Coherent-Decoding (WCD), a self-verifying test-time planning framework that treats WAM rollouts as falsifiable future--action hypotheses. At each decision step, WCD samples multiple candidates from a frozen WAM and ranks them using internal generative signals: flow-based video surprisal for visual plausibility and action path effort for action-generation stability. After execution, the realized observation audits the selected imagination, yielding an imagination--reality mismatch that trains a lightweight online predictor for future candidate selection. Thus, WCD converts delayed self-verification into pre-execution reliability estimation without updating the backbone model. On RoboTwin 2.0, WCD improves Hard success under limited randomized-scene supervision from $55.80\%$ to $60.90\%$, with a $+16.43$ gains on Horizon-3 tasks, and shows qualitative robustness on real Franka visual-shift tests. These results highlight a simple principle: test-time scaling for WAMs depends less on sampling more futures than on selecting reliable ones.