arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06508cs.RO

VLA-Corrector:面向基于提示的视觉-语言-动作策略闭环恢复的阶段感知可观测状态理解

VLA-Corrector: Stage-Aware Observable State Understanding for Prompt-Based Closed-Loop Recovery of Vision-Language-Action Policies

Chang Song, Bin Qian, Yan Feng, Zhijie Song

AI总结:

针对VLA策略长时程操作执行偏差问题,提出阶段感知验证与提示恢复框架,利用可观测历史习得验证器识别失败模式并生成恢复提示,在LIBERO基准上显著提升闭环可靠性。

AI中文摘要:

基于视觉-语言-动作(VLA)策略的长时程机器人操作在执行过程中仍易受偏差影响,因为最终任务成功与否几乎无法为诊断和纠正由动作噪声、物体位移或目标错位导致的失败提供信息。我们提出了一种阶段感知的失败验证与提示恢复框架,能够在无需参数更新或特权模拟器状态的情况下,实现对固定VLA策略的闭环纠正。该框架引入了一个基于可观测历史的习得验证器,通过对多视角视觉观测、本体感觉状态和执行动作进行时序建模,联合估计操作进度和执行风险。为了提供可解释的任务理解,我们将操作执行表示为语义进度阶段,包括接近、对齐、抓取、搬运和放置,并识别各阶段特有的失败模式。在检测到异常执行时,该框架保留原始指令并生成阶段条件的恢复提示,使同一个冻结的VLA策略能够产生纠正动作。在LIBERO和LIBERO Plus上的广泛多轮评估表明,所提方法在多种扰动下显著提高了闭环可靠性。在无法获取特权物体或目标坐标的情况下,习得验证器在评估设置中实现了接近特权规则验证器的恢复性能。这些结果表明,可观测的视觉-本体感觉-动作历史足以推断潜在任务状态,并为现有VLA策略实现实用的失败恢复。

英文摘要:

Long-horizon robot manipulation with Vision-Language-Action (VLA) policies remains vulnerable to execution-time deviations, as final task success provides little information for diagnosing and correcting failures caused by action noise, object displacement, or goal misalignment. We introduce a stage-aware failure verification and Prompt Recovery framework that enables closed-loop correction of a fixed VLA policy without parameter updates or privileged simulator states. The framework introduces an observable-history-based Learned Verifier that jointly estimates manipulation progress and execution risk by temporally modeling multi-view visual observations, proprioceptive states, and executed actions. To provide interpretable task understanding, we represent manipulation execution through semantic progress stages, including approach, alignment, grasp, transport, and placement, and identify stage-specific failure patterns. Upon detecting abnormal execution, the framework preserves the original instruction and generates a stage-conditioned recovery prompt, allowing the same frozen VLA policy to produce corrective actions. Extensive multi-round evaluations on LIBERO and LIBERO Plus demonstrate that the proposed approach substantially improves closed-loop reliability under diverse perturbations. Without access to privileged object or goal coordinates, the Learned Verifier achieves recovery performance close to that of the privileged rule-based verifier in the evaluated settings. These results show that observable visual-proprioceptive-action history is sufficient to infer latent task states and enable practical failure recovery for existing VLA policies.

↑