arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VLA-Scope:视觉-语言-动作模型的移位感知故障预测

VLA-Scope: Shift-Aware Failure Prediction for Vision-Language-Action Models

Kaiwen Zhu, Dongfang Liu, Liangkai Liu

arXiv 2609.21246首次发表:更新:

发表机构

Texas Tech University; Purdue University(德克萨斯理工大学; 普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

VLA-Scope提出两阶段框架,结合输入移位分类与执行历史,利用逻辑回归预测VLA模型在OOD条件下的执行失败,实验显示其预测性能优于基线。

AI 中文摘要

视觉-语言-动作(VLA)模型将视觉观察和自然语言指令映射为机器人动作,但分布移位可能损害其可靠性。由于这些模型在分布外(OOD)条件下仍可能成功,仅检测OOD输入不足以预测执行失败。本文提出VLA-Scope,一个两阶段框架,将输入移位特征与执行历史相结合,以预测OOD部署期间的失败。第一阶段使用池化的图像和语言表示来检测OOD输入并分类其移位类别。对于被标记为OOD的输入,第二阶段结合预测类别、动作前缀特征和执行进度特征。一个跨移位类别共享的逻辑回归模型在执行过程中更新失败风险。我们使用OpenVLA在十个LIBERO-Spatial任务上通过留一组交叉验证评估该框架。OOD检测的ROC-AUC达到0.9454,移位分类准确率达到91%。独立于OOD门控,对所有1,400次OOD部署进行评估,故障预测器在60个执行动作后的ROC-AUC为0.8497,而无需执行进度特征时为0.7906。其ROC-AUC也高于所评估的ActProbe和SAFE-MLP基线。这些结果表明,将动作特征与时间聚合的执行步骤表示相结合,可改善输入移位下的故障预测。

英文摘要

Vision-language-action (VLA) models map visual observations and natural-language instructions to robotic actions, but distribution shifts can compromise their reliability. Because these models may still succeed under out-of-distribution (OOD) conditions, detecting OOD inputs alone is insufficient to predict execution failure. In this paper, we introduce VLA-Scope, a two-stage framework that combines input-shift characterization with execution history to predict failure during OOD rollouts. The first stage uses pooled image and language representations to detect OOD inputs and classify their shift categories. For inputs flagged as OOD, the second stage combines the predicted category, action-prefix features, and execution progress features. A logistic regression model shared across shift categories updates failure risk as execution proceeds. We evaluate the framework with OpenVLA on ten LIBERO-Spatial tasks using leave-one-group-out cross-validation. OOD detection achieves a ROC-AUC of 0.9454, and shift classification achieves 91% accuracy. Evaluated independently of the OOD gate on all 1,400 OOD rollouts, the failure predictor achieves a ROC-AUC of 0.8497 after 60 executed actions, compared with 0.7906 without execution progress features. It also achieves a higher ROC-AUC than the evaluated ActProbe and SAFE-MLP baselines. These results suggest that combining action features with temporally aggregated execution step representations improves failure prediction under input shifts.

Comments9 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑