SAFECAST:结合对比集训练与校准的VLA策略鲁棒故障检测
SAFECAST: Robust Failure Detection for VLA Policies with Contrast-Set Training and Calibration
浏览论文内容
中文总结 AI 辅助
SAFECAST 借助对比集扰动优化隐状态探测器训练与校准,在真实和模拟实验中显著提升 VLA 策略故障检测 ROC-AUC,且视觉与语言对比集扰动结合时效果最优。
中文摘要 AI 辅助
视觉-语言-动作(VLA)策略在部署时的分布偏移(如杂乱环境、干扰对象、光照变化、新对象、初始状态改变、指令措辞修改)下常失效。基于隐状态的风险探测结合函数型共形预测可检测 rollout 故障,但其可靠性依赖校准数据与部署条件匹配。我们提出 SAFECAST,利用对比集扰动改进隐状态探测器的训练与部署时偏移校准。在真实世界 DROID 实验及多 VLM 主干的 LIBERO 模拟实验中,SAFECAST 相比 SOTA 基准的故障检测 ROC-AUC 分数有统计显著提升。进一步发现,同时使用视觉与语言对比集扰动增强数据时,SAFECAST 收益最大;且使用对比集扰动时,模拟到真实的校准仅用真实 rollout 数据即可得到更优探测器。
英文摘要
Vision-language-action policies often fail under deployment-time distribution shifts such as clutter, distractor objects, lighting changes, novel objects, altered initial states, and reworded instructions. Hidden-state-based risk probes combined with functional conformal prediction can detect rollout failures, but their reliability depends on calibration data matching deployment conditions. We introduce SAFECAST, which leverages contrast set perturbations to improve hidden-state probe training and calibration for deployment time shift. SAFECAST statistically significantly improves failure detection ROC-AUC scores over a state of the art baseline in both real-world DROID and LIBERO simulation experiments across multiple VLM backbones. We further find that SAFECAST benefits most when both visual and language contrast set perturbations are used to augment data, and that with contrast set perturbations, sim-to-real calibration leads to better probes than using real rollout data only.
发表机构
- University of Southern California(南加州大学)
机构由 AI 辅助整理,请以论文原文为准。