arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GUARD:面向扩散型视觉-语言动作(VLA)的不确定性与基于消融的风险检测

GUARD: Grounding Uncertainty and Ablation-Based Risk Detection for Diffusion-Based VLAs

Suhas Hegde, Jitendra Yasaswi Bharadwaj Katta

arXiv 2608.04510首次发表:更新:

AI 中文总结

本文提出GUARD方法,无需修改预训练扩散型VLA策略即可检测其故障,在多基准测试中提升未见过任务的ROC-AUC,提供可跨多维度迁移的故障信号。

AI 中文摘要

基于扩散的视觉-语言动作(VLA)策略即便在预测未充分锚定任务定义的视觉与语言证据时,仍能生成看似合理的动作。本文提出GUARD,一种测试时故障检测方法,无需修改预训练策略即可测量这种锚定程度。GUARD估算最终视觉-语言模型键值(KV)缓存中按token索引的条目的影响,通过消融显著KV条目构建反事实缓存,并将其去噪响应与原始条件进行比较。基于该比较,得到GUARD诊断流,包含敏感性、注意力熵、模态偏差和锚定效率,这些指标经在线校准后由轻量时序分类器处理。在五个策略基准设置的任务保留拆分下对GUARD进行评估,使用Pi0、SmolVLA和Alpamayo-1.5模型在LIBERO、SimplerEnv、MetaWorld和PhysicalAI-AV数据集上测试。GUARD在五个未见过任务设置中的四个上取得最佳ROC-AUC,在剩余设置中排名第二,相比最强的竞争运行时监控器,未见过任务的平均ROC-AUC提升5.73个百分点,且与最佳见过任务平均值的差距在0.19个百分点以内。这些结果表明,直接探测动作头对多模态证据的依赖,可提供跨策略、任务、实体和领域的可迁移故障信号。

英文摘要

Diffusion-based vision-language-action (VLA) policies can generate plausible actions even when their predictions are weakly grounded in the visual and language evidence defining the task. We introduce GUARD, a test-time failure detection method that measures this grounding without modifying the pretrained policy. GUARD estimates the influence of token-indexed entries in the final vision-language model key-value (KV) cache, constructs counterfactual caches by ablating salient KV entries, and compares their denoising responses with the original conditioning. Based on the comparison, we derive GUARD diagnostic stream including sensitivity, attention entropy, modality bias, and grounding efficiency, which are calibrated online and processed by a lightweight temporal classifier. We evaluate GUARD under task-held-out splits across five policy-benchmark settings, using Pi0, SmolVLA, and Alpamayo-1.5 on LIBERO, SimplerEnv, MetaWorld, and PhysicalAI-AV. GUARD achieves the best ROC-AUC on four of five unseen-task settings and ranks second on the remaining setting, improving the average unseen-task ROC-AUC by 5.73 percentage points over the strongest competing runtime monitor while remaining within 0.19 points of the best seen-task average. These results show that directly probing action-head dependence on multimodal evidence provides a transferable failure signal across policies, tasks, embodiments, and domains.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑