arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PerturBot:通过扰动训练打破视觉-语言-动作模型中的捷径先验

PerturBot: Breaking Shortcut Priors in Vision-Language-Action Models with Perturbative Training

Mingyu Liu, Chonghao Sima, Tianjian Feng, Hanqing Wang, Cong Chen, Hao Chen, Chunhua Shen

arXiv 2610.04616首次发表:更新:

发表机构

Zhejiang University; Shanghai Innovation Institute; University of Hong Kong; HKUST(GZ)(浙江大学; 上海创新研究院; 香港大学; 香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Perturbot训练框架与GroundingFscore评估指标,通过扰动和重标注打破VLA模型中的模态捷径,使其依赖任务证据而非捷径,实现健康扩展。

AI 中文摘要

视觉-语言-动作(VLA)策略能够完成复杂任务,却忽略本应决定其动作的证据。靠近腕部相机的物体可以偏移指令目标。语言和动作表现出相同模式:即使动词改变,熟悉的名词也能触发训练中与之配对的操作为;夹爪在空抓时也可能抬起。我们将这些依赖称为模态捷径:成功演示中的规律性使得视觉、词汇或运动线索足以预测专家动作,而无需底层决策所需的任务证据。更多同类演示可以提高任务成功率,但捷径依然存在。我们提出Perturbot,使任务相关证据更易使用,并让捷径单独不足以决策:它应用保持任务的腕部视角扰动,用决策相关字幕丰富指令,并添加随机和失败轨迹片段,以它们包含的行为重新标注。它通过改变缩放的内容来补充缩放,且不改变推理过程。此外,我们提出GroundingFscore,一种离线评分,用于诊断策略依赖模态捷径的严重程度。任务成功率显示策略是否改进,而GroundingFscore揭示策略是否健康扩展,即依赖任务证据而非捷径。Perturbot和GroundingFscore共同提供了一个训练与评估框架,用于将VLA决策与捷径先验分离,同时保持对任务相关证据的响应性。

英文摘要

A vision--language--action (VLA) policy can complete complex tasks while ignoring the evidence that should determine its actions. An object held near the wrist camera can displace the instructed target. Language and action show the same pattern: a familiar noun can trigger the operation it was paired with in training even after the verb changes, and a gripper that closed on nothing may lift anyway. We call these dependencies modality shortcuts: regularities in successful demonstrations make visual, lexical, or motor cues sufficient to predict expert actions without the task evidence needed for the underlying decision. More demonstrations of the same kind can raise task success while leaving these shortcuts intact. We propose Perturbot which makes task-relevant evidence easier to use and shortcuts insufficient on their own: it applies task-preserving wrist-view perturbations, enriches instructions with decision-relevant captions, and adds random and failed trajectory segments relabeled with the behavior they contain. It complements scaling by changing what is scaled, and leaves inference unchanged. Moreover, we propose GroundingFscore, an offline score that diagnoses how severely a policy relies on modality shortcuts. Task success rate shows whether a policy improves, while GroundingFscore reveals whether the policy scales healthily, relying on task evidence rather than shortcuts. Together, Perturbot and GroundingFscore provide a training-and-evaluation framework for disentangling VLA decisions from shortcut priors while preserving responsiveness to task-relevant evidence.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑