arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39971cs.ROcs.LG

当指令检索轨迹:诊断并缓解VLA模型中的泛化失败

When Instructions Retrieve Trajectories: Diagnosing and Mitigating Generalization Failures in VLA Models

Hung-Jen Chen, Yu-Hsun Hou, Yan-Hong Chen, Yan-Fu Chen, Binghua Cai, Min Sun, Chun-Yi Lee

首次发表
浏览论文内容

中文总结 AI 辅助

针对VLA模型在反事实变化下因指令-动作绑定失败而泛化不足的问题,提出等变反事实训练(ECT),通过配对演示与损失提升位置交换成功率,在LIBERO-PRO、CALVIN及真实UR5e上均显著改善。

中文摘要 AI 辅助

视觉-语言-动作(VLA)模型在分布内任务上可超过90%的成功率,并能承受保持所需动作不变的干扰变化,但在要求不同动作的反事实变化下却会失败。因此,聚合鲁棒性分数可能掩盖一种更具体的失败,即策略同时响应语言和视觉,却不将两者结合以选择任务所需的动作。我们将这种失败称为指令-动作绑定。指令提示熟悉的轨迹族,视觉反馈调整其执行。对微调后的π0.5和GR00T-N1.7策略的行为分析显示,失败的轨迹常常保留源行为或切换到另一个已演示的任务。这些切换表明语言并非被简单忽略。读出和干预将这些选择与任务条件化的内部状态联系起来。我们对模仿目标的分析表明,狭窄的条件动作支持如何使基于环境的和指令键控的解决方案在演示上难以区分。这激发了等变反事实训练(ECT),它在两个层面起作用。ECT数据提供有效的演示,其中相同指令在可区分的场景中要求不同的动作,而ECT损失在同一更新中训练每个演示及其对应物。在受控的LIBERO-PRO比较中,完整ECT将π0.5的平均位置交换成功率从36%提升到59%。在CALVIN上,对应物已存在于原始数据中,ECT损失在无新演示的情况下提升了五任务完成率。在固定演示预算下的真实UR5e上,完整ECT将未见位置的成功率从8%提升到88%。

英文摘要

Vision-language-action (VLA) models can exceed 90% success on in-distribution tasks and withstand nuisance changes that preserve the required action, yet fail under counterfactual changes that demand a different action. Aggregate robustness scores can therefore conceal a more specific failure, in which a policy responds to both language and vision yet does not combine them to select the action the task requires. We call this failure instruction-action binding. Instructions cue familiar trajectory families, and visual feedback adjusts their execution. Behavioral analyses of fine-tuned $π_{0.5}$ and GR00T-N1.7 policies reveal that failed rollouts often retain the source behavior or switch to another demonstrated task. These switches show that language is not simply ignored. Readouts and interventions connect these choices to task-conditioned internal states. Our analysis of the imitation objective shows how narrow conditional action support can leave grounded and instruction-keyed solutions indistinguishable on the demonstrations. This motivates Equivariant Counterfactual Training (ECT), which acts at two levels. ECT data supply valid demonstrations in which the same instruction requires different actions in distinguishable scenes, while the ECT loss trains each demonstration with its counterpart in the same update. In a controlled LIBERO-PRO comparison, full ECT raises $π_{0.5}$'s mean position-swap success from 36% to 59%. On CALVIN, where counterparts already occur in the original data, the ECT loss improves five-task completion without new demonstrations. On a real UR5e under a fixed demonstration budget, full ECT raises unseen-position success from 8% to 88%.

发表机构

  • National Tsing Hua University(国立清华大学)
  • National Taiwan University(国立台湾大学)

机构由 AI 辅助整理,请以论文原文为准。

↑