当未来一致时提交:面向机器人操作的结果感知自适应动作分块
Commit While Futures Agree: Consequence-Aware Adaptive Action Chunking for Robot Manipulation
另 2 家 · 查看机构详情
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
- ANU Intelligent(澳大利亚国立大学智能)
- Victoria University of Wellington(惠灵顿维多利亚大学)
- The Hong Kong University of Science and Technology(香港科技大学)
- Southern University of Science and Technology(南方科技大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
提出结果感知自适应动作分块(CA$^3$C),基于未来一致性决定提交或重规划,无需修改基础策略,显著降低机器人操作失败率。
中文摘要 AI 辅助
动作分块策略预测多步控制序列,但一个基本问题仍然存在:在重新规划之前,应提交预测动作分块中的多少内容?现有系统通常执行固定长度的前缀,隐含地假设相同的执行范围在不同状态下都保持可信。一些自适应方法根据预测动作的相似性或稳定性来估计该范围。然而,不同的动作可能导致相同成功的结果,而相似的动作可能产生不同的未来,这表明提交应由想象未来的一致性而非动作空间中的相似性来决定。为此,我们提出了结果感知自适应动作分块(CA$^3$C),这是一个基于简单原则的推理时框架:当想象未来一致时提交,当它们分歧时重新规划。在不修改或重新训练基础策略的情况下,CA$^3$C使用动作条件世界模型,在相同采样噪声下想象多个候选动作分块的未来结果。利用这些想象的结果,我们将执行范围估计表述为贝叶斯变点推断问题,并通过未来共识选择执行候选。在多个模拟基准和真实世界机器人操作任务中,CA$^3$C持续改进多样化的动作分块策略,相对于相应基础策略,失败率相对降低高达71.8%。
英文摘要
Action-chunking policies predict multi-step control sequences, but a fundamental question remains: how much of a predicted action chunk should be committed before replanning? Existing systems typically execute a fixed-length prefix, implicitly assuming that the same execution horizon remains trustworthy across states. Some adaptive methods estimate this horizon from the similarity or stability of predicted actions. However, different actions may lead to the same successful outcome, whereas similar actions can produce different futures, suggesting that commitment should be determined by agreement among imagined futures rather than by similarity in action space. To this end, we propose Consequence-Aware Adaptive Action Chunking (CA$^3$C), an inference-time framework built on a simple principle: commit while imagined futures agree, and replan when they diverge. Without modifying or retraining the base policy, CA$^3$C uses an action-conditioned world model to imagine the future consequences of multiple candidate action chunks under the same sampling noise. Using these imagined consequences, we formulate execution-horizon estimation as a Bayesian change-point inference problem and select the execution candidate through future consensus. Across multiple simulation benchmarks and real-world robot manipulation tasks, CA$^3$C consistently improves diverse action-chunking policies, achieving up to a 71.8% relative reduction in failure rate over the corresponding base policies.