arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

编码但不受控:揭示视觉-语言机器人策略中的接地缺口

Encoded but Not in Control: Revealing the Grounding Gap in Vision-Language Robot Policies

Shaohan Jiang, Jiahang Cao, Qiduo He, Fengting Deng, Kun Wu, Jingkai Sun, Jiaxu Wang, Qiang Zhang, Qihao Zheng, Chunfeng Song, Ping Luo, Andrew F. Luo

arXiv 2610.06235首次发表:更新:

发表机构

HKU; CUHK(SZ); Horizon Robotics; Shanghai AI Lab; CUHK; USTC(香港大学; 香港中文大学(深圳); 地平线机器人; 上海人工智能实验室; 香港中文大学; 中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过场景保持的指令干预揭示视觉-语言机器人策略存在接地缺口:任务成功不保证指令遵循,编码的指令未能控制动作,并提供了诊断框架和进展标准。

AI 中文摘要

指令遵循是语言条件化机器人策略的核心:当同一场景允许多个有效动作时,语言应决定执行什么。然而,仅凭成功执行无法确定策略是遵循指令还是从场景推断任务。我们通过保持场景不变的指令干预来研究这种模糊性,使用有效目标替换、任意名词和不相关句子。我们在仿真和真实世界实验中评估了视觉-语言-动作(VLA)策略和世界-动作模型(WAMs)。我们的分析回答了三个问题:(a)任务成功是否意味着指令遵循?当指令要求不同的可见对象时,所有评估的策略主要接近并拾取与场景相关的错误原始目标。(b)这种失败是否由语言不敏感性引起?指令扰动影响任务性能。分层动作透镜显示中间动作预测对这些扰动有响应。线性探针准确恢复被指示的目标,表明修改后的指令被编码,尽管很少决定目标选择。(c)为什么编码的语言未能控制动作?注意力分析表明指令标记对动作生成的贡献较弱。目标标记注意力可能保持对原始对象的关注,揭示了目标编码与视觉接地之间的不匹配。UMAP和共享非负矩阵分解表明,目标信息在日益由场景身份组织的表示中仍可访问。我们的发现暴露了被名义成功掩盖的接地缺口,并提供了一个诊断框架。它们进一步建立了进展的具体标准:策略应可靠地遵循用户意图的有效变化,即使这些变化与场景偏好的行为冲突。

英文摘要

Instruction following is central to language-conditioned robot policies: language should determine what to do when the same scene permits multiple valid actions. Yet successful execution alone cannot establish whether a policy follows the instruction or infers the task from the scene. We study this ambiguity through scene-preserving instruction interventions, using valid target substitutions, arbitrary nouns, and unrelated sentences while holding the scene fixed. We evaluate vision-language-action (VLA) policies and world-action models (WAMs) in simulation and in real-world experiments. Our analysis addresses three questions: (a) Does task success imply instruction following? When instructions request a different visible object, all evaluated policies predominantly approach and pick up the incorrect original target associated with the scene. (b) Is this failure caused by language insensitivity? Instruction perturbations affect task performance. A layerwise action lens shows intermediate action predictions respond to these perturbations. Linear probes accurately recover instructed targets, indicating modified instructions are encoded despite rarely determining target selection. (c) Why does encoded language fail to control action? Attention analysis indicates weak instruction-token contributions to action generation. Target-token attention can remain focused on the original object, revealing a mismatch between target encoding and visual grounding. UMAP and shared non-negative matrix factorization show target information remains accessible within representations increasingly organized by scene identity. Our findings expose a grounding gap concealed by nominal success and provide a diagnostic framework. They further establish a concrete criterion for progress: policies should reliably follow valid changes in user intent, even when they conflict with scene-favored behavior.

Comments30 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑