发表机构
University of Michigan; NVIDIA(密歇根大学; 英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对世界行动模型,提出无需训练的选择性测试时缩放框架\methodgated,基于预测未来的跨视图深度重投影一致性排序,经实验验证该方法能提高任务成功率,同时减少额外采样决策点,还识别出相关失败模式。
AI 中文摘要
测试时缩放通过额外计算改进基础模型推理,但机器人控制需在执行动作前决定额外计算是否有用。世界行动模型(WAMs)使此决策自然化。我们提出了\methodgated,一种用于WAMs的无需训练的选择性测试时缩放框架。首先实例化\method,一个固定预算的最佳N选择器,通过预测未来的跨视图深度重投影一致性对采样的展开进行排序。\methodgated添加了一个轻量级的动作-未来一致性门,仅在初始展开内部不一致时调用\method。在五个基准设置上的实验表明,固定预算的\method在每个设置中都提高了N=8任务的成功率,启用门控后,\methodgated平均恢复了74.8%的始终开启的成功率提升,同时仅在26.2%的决策点触发额外采样。离线诊断表明跨视图重投影是一个强大的无任务标签选择器,我们将错误的低分选择识别为一种失败模式,有助于解释为什么随着N的增加性能会饱和或下降。
英文摘要
Test-time scaling improves foundation-model inference by spending additional computation, but robot control requires deciding whether extra compute is useful before executing an action. World Action Models (WAMs) make this decision natural: each rollout exposes both an action chunk and predicted future observations. We propose \methodgated, a training-free selective test-time scaling framework for WAMs. We first instantiate \method, a fixed-budget Best-of-$N$ selector that ranks sampled rollouts by cross-view depth reprojection consistency of their predicted futures, computed with a frozen geometry foundation model. \methodgated\ adds a lightweight action--future consistency gate that invokes \method\ only when the initial rollout appears internally inconsistent. Across five benchmark--backbone settings on RoboCasa, LIBERO Long, and RoboTwin~2.0, fixed-budget \method\ improves $N{=}8$ task success in every setting, e.g., raising the RoboCasa group average from $66.3\%$ to $68.4\%$ with Cosmos Policy and from $80.8\%$ to $82.5\%$ with X-WAM. With gating enabled, \methodgated\ recovers on average $74.8\%$ of the always-on success gain while triggering additional sampling on only $26.2\%$ of decision points. Offline diagnostics show that cross-view reprojection is a strong task-label-free selector, and we identify false low-score selections as a failure mode that helps explain why performance can saturate or degrade as $N$ increases.
CommentsExtened version of CVPR 2026 EAI workshop