arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

继续还是重规划?用于自适应执行时长的伯努利继续策略学习

Continue or Replan? Bernoulli-Continuation Policy Learning for Adaptive Horizon Execution

Weichen Xu, Zhenhua Liu, Lin Luo, Yaobo Liang, Chengtang Yao, Qingyu Mei, Jian Cao, Xixin Cao, Xing Zhang, Jiaolong Yang, Baining Guo

arXiv 2608.03483首次发表:更新:

AI 中文总结

针对现有VLA模型固定执行时长重规划的局限,提出BCP框架,通过强化学习训练自适应时长选择,在多个仿真任务、真实机器人任务上提升了成功率,且运行时间更短。

AI 中文摘要

现有的基于分块的视觉-语言-动作(Vision-Language-Action, VLA)模型在重规划前会执行固定数量的动作(即执行时长),将重规划变成与任务进度无关的、与任务无关的周期性调度。因此,当关键操作阶段前没有重规划边界时,该阶段会从过时的分块执行,而非从新重规划的分块执行。为解决这一局限,我们提出伯努利继续策略(Bernoulli-Continuation Policy, BCP),这是一种轻量、即插即用的自适应执行时长框架,可保持基础VLA模型冻结。给定固定长度的动作分块,其继续头将执行时长选择分解为一系列继续或重规划的决策,这对候选执行时长施加了有序的、前缀共享的归纳偏置,而非将其视为独立类别。由于每个分块的最优执行时长不可观测,我们使用轨迹级结果的强化学习训练该头,并引入重规划效率奖励,同时奖励任务成功和VLA的高效使用,防止策略坍缩为不必要的短执行时长。在RoboTwin 2.0上,以LingBot-VLA为基础策略,BCCP在13个低成功率任务上将平均成功率提高了+11.08%,在全部50个任务上从89.88%提高到93.94%(+4.06%)。尽管仅在干净(Clean)设置下训练,BCP可泛化到随机(Randomized)设置,将平均成功率提高+4.06%。它还可迁移到不同的基础策略π₀.₅,在LIBERO上取得更好结果(+1.7%),尤其在更难的LIBERO-PRO上提高+6.8%。在真实机器人上,BCP在两个操作任务上将成功率从74%提升至92%,从44%提升至84%。同时,其可忽略的开销结合更高的成功率,使BCP的总运行时间甚至低于固定执行时长的基线。

英文摘要

Existing chunk-based Vision-Language-Action (VLA) models execute a fixed number of actions (i.e., execution horizon) before replanning, turning replanning into a task-agnostic periodic schedule that is independent of task progress. As a result, when no replanning boundary falls before a critical manipulation stage, it is executed from a stale chunk rather than a freshly replanned one. To address this limitation, we propose Bernoulli-Continuation Policy (BCP), a lightweight, plug-and-play framework for adaptive horizon execution that keeps the base VLA frozen. Given a fixed-length action chunk, its continuation head decomposes execution-horizon selection into a sequence of continue-or-replan decisions, which imposes an ordinal, prefix-sharing inductive bias over candidate horizons rather than treating them as independent classes. Since the optimal horizon for each chunk is not observable, we train this head with reinforcement learning from trajectory-level outcomes and introduce a Replanning-Efficiency Reward that jointly rewards task success and efficient VLA usage, discouraging the policy from collapsing to unnecessarily short horizons. On RoboTwin 2.0 with LingBot-VLA as the base policy, BCP improves the average success rate by +11.08% on 13 low-success tasks and from 89.88% to 93.94% (+4.06%) across all 50 tasks. Although trained only under the Clean setting, BCP generalizes to the Randomized setting, raising the average success rate by +4.06%. It also transfers to a different base policy $π_{0.5}$, achieving a better result on LIBERO (+1.7%) and, notably, on the harder LIBERO-PRO (+6.8%). On a real robot, BCP lifts success from 74% to 92% and from 44% to 84% on two manipulation tasks. Meanwhile, its negligible overhead, combined with higher success, makes BCP's overall runtime even lower than the fixed-horizon baselines.

CommentsProject page: https://fleetfootwork.github.io/BCP/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑