快速规划,忠实执行:缩小分层视觉-语言-动作模型中的规划-执行差距
Fast Plans, Faithful Actions: Closing the Planning-Execution Gap in Hierarchical Vision-Language-Action Models
浏览论文内容
中文总结 AI 辅助
针对分层VLA模型中规划速度慢且计划未被充分利用的问题,提出Block-AR解码和归一化目标调制,将VLM前向次数从57降至8,延迟降低8.7倍,成功率显著提升。
中文摘要 AI 辅助
分层视觉-语言-动作(VLA)系统由高层视觉-语言规划器和生成连续动作的低层动作专家组成。这种分层设计只有在规划器能够足够快地生成计划以满足实时控制要求,并且生成的计划确实有助于动作生成时,才具有实用价值。我们研究了这样一个系统,即从$\pi_{0.5}$改编的航点分层流水线,并发现这两个要求均未得到满足。该基线依赖于令牌级自回归解码(Token-AR)来生成航点计划,需要57次非常昂贵的视觉-语言模型(VLM)前向传播。然而,我们发现擦除航点端点对任务成功率几乎没有影响。这两个发现揭示了规划器与执行器之间的错位:规划器以过细的粒度生成输出,而执行器未充分利用计划作为控制条件。我们通过航点对齐的块自回归解码(Block-AR)解决延迟问题,并通过归一化目标调制(NGM)解决计划利用不足的问题,这是一种受相位门控和防捷径训练约束的逐层目标路径,使得航点影响动作生成同时保留其他信号。我们的方法在LIBERO上将VLM前向传播的最大次数从57次减少到8次(包括一次前缀预填充),并在Rokae双臂机器人上实现了$8.7\times$的规划延迟降低。通过归一化目标调制和防捷径训练,Block-AR在LIBERO-Long上的成功率从91.0%提升至96.2%,在四个套件上的平均成功率从95.85%提升至98.45%。在该机器人的三个双臂任务上,各方法的成功率保持相当。
英文摘要
Hierarchical vision-language-action (VLA) systems consist of a high-level vision-language planner and a low-level action expert that generates continuous actions. This hierarchical design has practical value only if the planner can generate plans fast enough to meet real-time control requirements, and the resulting plans actually contribute to the generation of action. We study one such system, a waypoint hierarchy pipeline adapted from $π_{0.5}$, and find that neither requirement is satisfied. This baseline relies on token-level autoregressive decoding (Token-AR) to generate a waypoint plan, requiring 57 very expensive vision-language model (VLM) forward passes. However, we find that erasing the waypoint endpoints has little effect on task success. Two findings reveal the misalignment of planner-executor: the planner generates outputs at an excessively fine granularity, and the executor underuses plans as a control condition. We address the latency issue with waypoint-aligned block-autoregressive decoding (Block-AR), and plan underuse issue with normalized goal modulation (NGM), a layer-wise goal path constrained by phase gating and anti-shortcut training so that the waypoint influences action generation maintaining other signals. Our method reduces the maximum number of VLM forward passes from 57 to 8 on LIBERO, including one prefix prefill, and achieves an $8.7\times$ reduction in planning latency on a Rokae dual-arm robot. With normalized goal modulation and anti-shortcut training, Block-AR's success rate on LIBERO-Long increases from 91.0% to 96.2%, while its average success rate across the four suites increases from 95.85% to 98.45%. On three bimanual tasks with this robot, success rates remain comparable across methods.
发表机构
- MARS Lab, Nanyang Technological University(南洋理工大学MARS实验室)
机构由 AI 辅助整理,请以论文原文为准。