arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10484cs.ROcs.AIcs.CL

在重建中迷失:将视觉-语言-动作模型中的动作表示与语言对齐

Lost in Reconstruction: Aligning Action Representations with Language in Vision-Language-Action Models

Li Wenjie, Yash Jangir, Ignacy Stepka, Yash Agarwal, Marion Kipsang, Yonatan Bisk

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对视觉-语言-动作模型中动作表示与语言脱节的问题,提出SALT语义对齐动作标记器,在SimplerEnv中使策略平均成功率达71.9%,显著优于基线方法。

中文摘要 AI 辅助

动作动词不仅描述动作的物理结果,还描述动作的执行方式。然而,视觉-语言-动作模型(VLAs)中的动作表示通常针对原始动作空间中的L1/L2损失下的重建进行优化,其中数值邻近性不一定反映语言上有意义的区分。在BridgeV2上,我们表明动作轨迹包含超越视觉状态变化的动词接地信息,且仅用于重建的离散标记化会系统性地侵蚀该信息。为解决此问题,我们引入SALT,即语义对齐动作标记器,它在VQ-VAE式标记器的基础上增加了辅助目标,要求冻结的视觉-语言模型从量化动作潜变量中恢复该片段指令。使用SALT训练的策略在SimplerEnv中实现了71.9%的平均成功率,而仅用重建的VQ-VAE标记器为42.7%,FAST为31.2%。SALT还开发了动词专用代码,同时保持重建保真度。这些结果表明,机器人动作轨迹提供了语言接地的来源,且在动作表示中保留此结构可显著改善语言条件控制。

英文摘要

Action verbs describe not only the physical outcomes of actions, but also how those actions are performed. Yet action representations in vision-language-action models (VLAs) are typically optimized for reconstruction under L1/L2 losses in raw action space, where numerical proximity need not reflect linguistically meaningful distinctions. On BridgeV2, we show that action trajectories contain verb-grounding information beyond visual state changes, and that reconstruction-only discrete tokenization systematically erodes this information. To address this problem, we introduce SALT, a Semantically ALigned action Tokenizer that augments a VQ-VAE-style tokenizer with an auxiliary objective requiring a frozen vision-language model to recover the episode instruction from quantized action latents. Policies trained with SALT achieve 71.9% average success in SimplerEnv, compared with 42.7% for a reconstruction-only VQ-VAE tokenizer and 31.2% for FAST. SALT also develops verb-specialized codes while maintaining reconstruction fidelity. These results show that robot action trajectories provide a source of language grounding and that preserving this structure in action representations can substantially improve language-conditioned control.

发表机构

  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

↑