arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当聆听变得更简单:为无捷径的视觉-语言-动作模型清除视觉线索

When Listening Becomes Easier: Scrubbing Visual Cues for Shortcut-Free VLAs

Jasper Gerigk, Kenzo Aspuru-Takata, Chin-Hsuan Wu, Mohammad Mohammadi, Shuhong Zheng, Igor Gilitschenski

arXiv 2610.10912首次发表:更新:

发表机构

University of Toronto; Vector Institute(多伦多大学; 向量研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对机器人学习中视觉捷径学习的问题,提出任务清除的领域对抗训练方法,可提升视觉-语言-动作模型的分布外鲁棒性并消除视觉捷径学习。

AI 中文摘要

捷径学习是机器人学习中普遍存在的问题。机器人演示数据集的多样性有限,可能会误导策略利用任务与无关特征(如视角或背景)之间的虚假相关性。收集足够多样的机器人演示成本高昂且效率低下,因此需要算法替代方案。我们聚焦于视觉-语言-动作(VLA)模型,发现不同的视觉-语言模型主干对视觉捷径学习的易感性存在显著差异。我们发现模型行为与我们提出的表示层面指标“动作间隔”相关,该指标无需策略回滚。视觉捷径始终会进入早期层的动作表示,不同模型在后期层通过整合语言信息纠正这些捷径的程度有所不同。为提升模型对语言的关注度,我们引入了任务清除(task scrubbing),这是一种新的领域对抗训练方法,可降低模型使用视觉捷径的可能性并提升VLA的泛化能力。在模拟环境和真实世界中针对多个VLA及视觉线索开展的实验表明,任务清除提升了分布外鲁棒性,且常能消除视觉捷径学习。

英文摘要

Shortcut learning is a prevalent issue in robot learning. The limited diversity of robot demonstration datasets can mislead policies into exploiting spurious correlations between tasks and irrelevant features, such as viewpoint or background. Collecting sufficiently diverse robot demonstrations is costly and inefficient, motivating algorithmic alternatives. We focus on vision-language-action (VLA) models and discover that different vision-language model backbones exhibit substantially different levels of susceptibility to visual shortcut learning. We find that model behavior correlates with our proposed representation-level metric, action margin, which requires no policy rollouts. Visual shortcuts consistently enter action representations in early layers, with models differing in the extent to which later layers correct them by incorporating language information. To boost models' attention to language, we introduce task scrubbing, a new domain-adversarial training method that decreases models' likelihood of using visual shortcuts and improves VLAs' generalization. Experiments in both simulation and the real world across multiple VLAs and visual cues show that task scrubbing improves out-of-distribution robustness and often eliminates visual shortcut learning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑