arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35439cs.ROcs.CV

修订而非重启:用于闭环世界-动作模型的可修订视觉规划

Revision, Not Restart: Revisable Visual Plans for Closed-Loop World-Action Models

Pengyiang Liu, Junbo Niu, Wenhao Zheng, Xinchen Chen, Canyu Li, Zhongyue Shi, Jiahao Xie, Si Liu

首次发表
浏览论文内容

中文总结 AI 辅助

提出可修订时间规划(RTP),通过修订桥和自适应策略在反馈后修订视觉未来,提升闭环任务成功率,在RoboMME和RMBench上分别达48.6%和84.8%。

中文摘要 AI 辅助

世界-动作模型利用预测的视觉未来来调节机器人动作,然而执行反馈可能使预测的部分内容失效,同时其任务结构仍然有用。我们提出了可修订时间规划(RTP),它将视觉未来保持为持久的动作条件,并在反馈后对其进行修订。其核心机制是一个学习到的修订桥:它恢复视觉生成过程中保存的中间状态,并根据当前观测调整其后续生成。视觉和动作监督将这种修订与后续控制联系起来。时间感知的历史提供观测证据,自适应策略在解码下一个动作之前选择保留、桥接修订或从新噪声重新规划。在RoboMME和RMBench上,RTP分别实现了48.6%和84.8%的任务平均成功率。匹配比较支持学习到的延续;估计的检查点来源和动作前缀效应为正,但精度较低。这些结果将反馈驱动的视觉规划修订与闭环任务性能联系起来。项目页面:此https URL

英文摘要

World-action models use predicted visual futures to condition robot actions, yet execution feedback can invalidate parts of a prediction while leaving its task structure useful. We propose Revisable Temporal Planning (RTP), which maintains the visual future as a persistent action condition and revises it after feedback. Its central mechanism is a learned revision bridge: it resumes an intermediate state saved during visual generation and adapts its continuation to current observations. Visual and action supervision connect this revision to subsequent control. Time-aware history supplies observed evidence, and an adaptive policy selects retention, bridge revision, or fresh replanning from new noise before decoding the next action. On RoboMME and RMBench, RTP achieves task-averaged success rates of 48.6% and 84.8%, respectively. Matched comparisons support learned continuation; estimated checkpoint-source and action-prefix effects are positive but less precisely resolved. These results connect feedback-driven visual-plan revision to closed-loop task performance. Project Page: https://PLACEHOLDER.github.io/RTP/

发表机构

  • Beihang University(北京航空航天大学)
  • Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑