arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.26789cs.RO

CheckVLA:面向长 horizon 移动操作的基于动作条件世界模型的执行时验证

CheckVLA: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation

Yushan Liu, Peibo Sun, Xintao Chao, Zhenyang Yang, Yifan Xie, Lingfeng Zhang, Shoujie Li, Chenyu Tang, Fang Chen, Xiao-Ping Zhang, Wenbo Ding

首次发表
浏览论文内容

中文总结 AI 辅助

CheckVLA 是基于动作条件世界模型的长 horizon 移动操作执行时验证方法,在 RoboCasa365 上其平均成功率比定期重规划提升 8.5 个百分点,且及时召回率显著优于仅观测及动作打乱的对照方法。

中文摘要 AI 辅助

视觉-语言-动作(VLA)策略通常通过开环动作块执行长 horizon 移动操作,即发出多个动作而不接收新的高层视觉输入。因此,一个已提交的动作块隐含了观测应如何演变,但意外的偏差会违反这一预期,而剩余动作会继续传播错误:提交时的策略置信度无法对调度后发生的偏差做出反应,仅基于观测的异常分数缺乏基于动作条件的参考,以区分预期效果与无法解释的变化。我们提出 CheckVLA,它使用单独训练的冻结动作条件世界模型来验证执行过程。一个经保形校准的风险阈值限定了不必要首次干预的 episode 级概率,并确定何时进行干预;该阈值的超出程度控制重写后缀对被取代动作块的保留强度;感知延迟的硬前缀将替换限制在仍可部署的动作范围内;事件驱动的关键帧库则保留修复过程中先前进展的证据。在 RoboCasa365 上,采用通用训练方案和匹配的调用预算,CheckVLA 的平均成功率达 36.1%,而定期重规划的成功率为 27.6%(提升 8.5 个百分点)。在匹配的 5% episode 级误报目标下,动作条件将及时召回率提升至 77.9%,而仅观测控制的召回率为 48.6%,动作打乱控制的召回率为 37.9%。这些模拟结果表明,基于动作条件的验证可在块执行期间恢复反馈,同时保持修复与推理延迟的一致性。

英文摘要

Vision-language-action (VLA) policies commonly execute long-horizon mobile manipulation through open-loop action chunks, issuing multiple actions without receiving new high-level visual input. A committed chunk therefore implies how observations should evolve, but accidental deviations can violate this expectation while the remaining actions continue to propagate the error: commit-time policy confidence cannot react to a deviation that occurs after dispatch, and observation-only anomaly scores lack an action-conditioned reference for separating expected effects from unexplained changes. We propose CheckVLA, which verifies execution with a separately trained, frozen action-conditioned world model. A conformally calibrated risk threshold bounds the episode-level probability of an unnecessary first intervention and determines when to intervene, its exceedance controls how strongly the rewritten suffix retains the superseded chunk, latency-aware hard prefixing restricts replacement to actions that remain deployable, and an event-driven keyframe bank preserves evidence of prior progress across repairs. On RoboCasa365, under a common training recipe and a matched invocation budget, CheckVLA attains a 36.1% average success rate against 27.6% for periodic replanning (+8.5 points). At a matched 5% episode-level false-alarm target, action conditioning raises timely recall to 77.9%, against 48.6% for an observation-only control and 37.9% for an action-shuffled control. These simulation results support action-conditioned verification as a way to restore feedback during chunked execution while keeping the repair consistent with inference latency.

↑