发表机构
University of Bremen; University of North Texas; Toyota Motor North America(不来梅大学; 北得克萨斯大学; 丰田北美汽车公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对机器人操作中生成未来的时序错位问题,提出RAFC方法,结合FEC,在CALVIN和Franka上大幅提升时序不匹配下的任务成功率,增益达7.0个百分点。
AI 中文摘要
机器人即将执行的任务生成视频仅在描绘机器人实际所处阶段时才是有用的指导。本文表明,时序错位会使与任务一致的生成未来变为主动有害的指导。在CALVIN数据集上,5帧的提前偏移几乎消除了生成未来的益处,使成功率从无未来时的54.0%提升至81.3%,而施加的时序偏移则进一步将成功率降至34.2%,比无未来策略低19.8个百分点。本文提出了可靠性感知未来条件控制(RAFC),将该问题视为控制问题而非生成问题。RAFC在每一步都会估计对接收片段的信任程度以及应偏好哪个附近的时序假设,当两者都不匹配时则回归静态分支,且仅从任务奖励中学习,无需偏移标签或对齐监督。RAFC构建于未来经验条件控制(FEC)之上,FEC通过任务 grounding、无机器人数字孪生回退和无掩码视频扩散一次性构建片段。在故意的非网格阶段偏移和速率不匹配情况下,RAFC大幅提高了时序不匹配下的成功率。候选集成在对齐附近的恢复中占大部分,而学习到的可靠性在非网格偏移下,比相同候选库的均匀平均进一步提升了7.0个百分点。该增益在评估任务集上保持有效,并在无人为施加的自然时序不匹配下的Franka机器人上得到验证,总成功率从26.7%提升至56.7%。所有资源将公开提供。
英文摘要
A generated video of a task the robot is about to perform is useful guidance only if it depicts the phase the robot is actually in. We show that temporal misalignment can turn a task-consistent generated future into actively harmful guidance. On CALVIN, a five-frame early shift nearly erases the benefit of generated futures, reducing success from 81.3% to 54.8% against 54.0% without futures; imposed timing shifts reduce it even further to 34.2%, 19.8 points below the future-free policy. We introduce Reliability-Aware Future Conditioning (RAFC), which treats this as a control problem rather than a generation problem. At every step, RAFC estimates how far to trust the received clip and which nearby temporal hypothesis to prefer, falling back toward a static branch when neither fits, and it learns both from task reward alone without shift labels or alignment supervision. RAFC sits on top of Future-Experience Conditioning (FEC), which builds the clip once from task grounding, a robot-free digital-twin rollout, and mask-free video diffusion. Under deliberately off-grid phase shifts and rate mismatch, RAFC substantially improves success under temporal mismatch. Candidate ensembling accounts for most of the recovery near alignment, while learned reliability adds a further 7.0 percentage points over uniform averaging of the identical candidate bank under off-grid shifts. The gain holds on the evaluated task sets and survives on a Franka under natural timing mismatch nobody imposed, where aggregate success rises from 26.7% to 56.7%. All resources will be made publicly available. https://future-condition.github.io/.