arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

无需训练的、基于语言反馈的VLA模型错误动作校正方法

Training-Free Action Correction for VLA Model Failures via Language Feedback

Owen Kwon, Pablo Ortega-Kral, Arthur Bucker, Jean Oh

arXiv 2608.29967首次发表:更新:

发表机构

Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出CorrectVLA框架,利用任务级自然语言反馈在不重训的情况下校正VLA模型的执行错位错误,在仿真与真实机器人实验中均实现了良好的泛化校正效果,明确了推理阶段动作校正的适用边界。

AI 中文摘要

视觉-语言-动作(Vision-Language-Action, VLA)模型具备较强的语义理解能力,但在部署过程中会出现系统性错误。这类错误的发生条件,以及能否在不进行重新训练的情况下对其进行校正,目前仍未得到充分理解。本文针对这一空白开展研究,提出CorrectVLA框架,该框架将任务级自然语言校正转化为动作幅度的附加调整,且不修改策略权重。仅需提供一次任务级校正,即可均匀应用于所有滚动执行过程,无需针对每个回合进行干预。在仿真实验中,CorrectVLA可校正分布内(in-distribution)及分布外(OOD)任务的执行错位错误。在UFactory xArm7机器人的真实实验中,当环境发生变化时,基础策略几乎完全失效,而CorrectVLA可恢复接近完美的成功率,且能泛化到不同的物体位置与身份。通过对LIBERO-90数据集的错误模式分类,研究发现:执行错位错误(策略到达正确目标但动作幅度校准错误)属于可校正子集,而语义理解本身失效的其他错误模式则不适用于该方法。该方法在策略具备策略正确性时有效,在缺乏基础理解时失效,为推理阶段的校正建立了实用的操作边界。

英文摘要

Vision-Language-Action (VLA) models demonstrate strong semantic understanding yet exhibit systematic failures during deployment. The conditions under which these failures occur, and whether they can be corrected without retraining, remain poorly understood. In this paper, we take steps toward addressing this gap. We present CorrectVLA, a framework that translates task-level natural language corrections into additive action magnitude adjustments without modifying policy weights. A human provides a single task-level correction, applied uniformly across all rollouts without per-episode intervention. In simulation, CorrectVLA recovers execution misalignment failures across both in-distribution and OOD tasks. In real-robot experiments on a UFactory xArm7 under environment shift, CorrectVLA restores near-perfect success where the base policy almost entirely breaks down, generalizing across object locations and identities. Through a taxonomy of failure modes on LIBERO-90, we find that execution misalignment failures, where the policy reaches the correct target but miscalibrates action magnitudes, represent the correctable subset, while other failure modes where semantic comprehension itself breaks down are not amenable to this approach. The approach succeeds when policies possess strategic correctness and fails when fundamental comprehension is absent, establishing a practical operational boundary for inference-time correction.

Comments8 pages, 6 figures. Project page: https://correctvla.github.io

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑