FailureSpot:面向视觉-语言-动作模型的标签高效时间戳级故障检测
FailureSpot: Label-Efficient Timestamp-Level Failure Detection for Vision-Language-Action Models
查看机构详情
- Wayne State University(韦恩州立大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对视觉-语言-动作模型的长 horizon 故障检测痛点,提出数据高效框架,结合弱监督与主动学习提升时间戳级和轨迹级故障检测性能。
中文摘要 AI 辅助
视觉-语言-动作(VLA)策略在通用机器人操纵中展现出强大潜力,但在长 horizon 执行过程中仍可能出现不可预测的故障,因此可靠的故障检测对安全部署至关重要。现有方法要么依赖视觉模型,通常仅在错误动作发生后才能检测到故障;要么使用在 VLA 内部表征上训练的轻量型主动检测器。然而,这些主动方法通常以轨迹级标签进行监督,导致失败轨迹中正常的故障前行为被错误标记为故障,这种监督不匹配会引入标签噪声,并限制轨迹级检测精度和精确的时间戳级故障定位。本研究针对细粒度时间戳级 VLA 故障检测展开,同时解决密集标注的成本问题。我们提出一种数据高效框架:首先利用未标注的 VLA 动作块构建动作衍生的弱监督信号,捕捉诸如连续块不一致、冻结或空闲动作、激进随机运动等异常模式;随后采用主动学习仅选择最不确定的轨迹进行时间戳级标注,并使用这些信息丰富的标签对检测器进行微调。在多个 VLA 策略上开展的实验表明,我们的方法可同时提升时间戳级和轨迹级故障检测性能。
英文摘要
Vision-language-action (VLA) policies have shown strong potential for general-purpose robotic manipulation, but they can still fail unpredictably during long-horizon execution, making reliable failure detection essential for safe deployment. Existing methods either rely on visual models that typically detect failures only after erroneous actions have occurred, or use lightweight proactive detectors trained on VLA internal representations. However, these proactive methods are often supervised with trajectory-level labels, causing normal pre-failure behavior in unsuccessful trajectories to be incorrectly labeled as failure. This supervision mismatch introduces label noise and limits both trajectory-level detection accuracy and precise timestamp-level failure localization. In this work, we study fine-grained timestamp-level VLA failure detection while addressing the cost of dense annotation. We propose a data-efficient framework that first leverages unlabeled VLA action chunks to construct action-derived weak supervision signals, capturing abnormal patterns such as inconsistent consecutive chunks, frozen or idle actions, and aggressive random motions. We then use active learning to select only the most uncertain trajectories for timestamp-level annotation and fine-tune the detector with these informative labels. Experiments across multiple VLA policies show that our method improves both timestamp-level and trajectory-level failure detection performance.