arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

回望以向前:生成式机器人策略的时间验证

Looking Back to Move Forward: Temporal Verification for Generative Robot Policies

Haoxuan Wang, Wayne Wu, Yan Yan, Bolei Zhou

arXiv 2609.39038首次发表:更新:

发表机构

University of Illinois Chicago; University of California, Los Angeles(伊利诺伊大学芝加哥分校; 加州大学洛杉矶分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出时间验证(TeV)框架,通过时间令牌和能量验证器对生成式机器人策略的动作块进行时间一致排序与引导,提升任务成功率并产生更平滑轨迹。

AI 中文摘要

生成式策略已成为机器人学习的一种有前景的范式,它将表达性生成动作建模与从大规模示范语料库进行的可扩展模仿学习相结合。然而,异质示范可能引发次优动作块,其误差随时间累积,最终将机器人推向难以恢复的分布外状态。动作验证提供了一种测试时扩展策略,通过采样多个候选动作并使用验证器选择一个执行来缓解这一失效模式。然而,现有方法在时间上是短视的且训练成本高昂,它们仅从当前观测评估候选动作,不考虑轨迹连续性,且常常依赖大型验证器和额外的专家示范。在本文中,我们提出了时间验证(TeV),一种用于流匹配VLA的高效时间感知动作验证框架。TeV首先学习一个时间令牌,总结最近的观测-动作历史,使候选块能够作为执行轨迹的延续而非孤立预测进行评估。基于该令牌,TeV构建正负对,无需额外专家示范或偏好标注,并对比训练一个基于能量的验证器,为更高质量、轨迹一致的动作块分配更低能量。除事后排序外,TeV还利用学习到的能量景观引导中间流样本朝向低能量区域,在最终选择前改进候选。在仿真和真实世界中的大量实验表明,TeV能够可靠地对动作候选进行排序,提高任务成功率,并产生更平滑的执行轨迹。

英文摘要

Generative policies have emerged as a promising paradigm for robot learning, combining expressive generative action modeling with scalable imitation learning from large demonstration corpora. However, heterogeneous demonstrations can induce suboptimal action chunks whose errors compound over time, eventually driving the robot into out-of-distribution states from which recovery is difficult. Action verification offers a test-time scaling strategy for mitigating this failure mode by sampling multiple candidate actions and using a verifier to select one for execution. Existing approaches, however, remain temporally myopic and costly to train, evaluating candidates from the current observation alone without accounting for trajectory continuity and often relying on large verifiers and additional expert demonstrations. In this paper, we introduce Temporal Verification (TeV), an efficient temporally aware action verification framework for flow-matching VLAs. TeV first learns a temporal token that summarizes recent observation--action history, enabling candidate chunks to be evaluated as continuations of the execution trajectory rather than as isolated predictions. Conditioned on this token, TeV constructs positive--negative pairs without additional expert demonstrations or preference annotations and trains an energy-based verifier contrastively to assign lower energy to higher-quality, trajectory-consistent action chunks. Beyond post-hoc ranking, TeV further uses the learned energy landscape to guide intermediate flow samples toward lower-energy regions, improving candidates before final selection. Extensive experiments in simulation and real-world settings demonstrate that TeV provides reliably ranks action candidates, improves task success rates, and produces smoother execution trajectories.

CommentsProject page at https://hatchetproject.github.io/tev/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑