在重建中迷失:将视觉-语言-动作模型中的动作表示与语言对齐
Lost in Reconstruction: Aligning Action Representations with Language in Vision-Language-Action Models
浏览论文内容
中文总结 AI 辅助
该研究针对视觉-语言-动作模型中动作表示与语言脱节的问题,提出SALT语义对齐动作标记器,在SimplerEnv中使策略平均成功率达71.9%,显著优于基线方法。
中文摘要 AI 辅助
动作动词不仅描述动作的物理结果,还描述动作的执行方式。然而,视觉-语言-动作模型(VLAs)中的动作表示通常针对原始动作空间中的L1/L2损失下的重建进行优化,其中数值邻近性不一定反映语言上有意义的区分。在BridgeV2上,我们表明动作轨迹包含超越视觉状态变化的动词接地信息,且仅用于重建的离散标记化会系统性地侵蚀该信息。为解决此问题,我们引入SALT,即语义对齐动作标记器,它在VQ-VAE式标记器的基础上增加了辅助目标,要求冻结的视觉-语言模型从量化动作潜变量中恢复该片段指令。使用SALT训练的策略在SimplerEnv中实现了71.9%的平均成功率,而仅用重建的VQ-VAE标记器为42.7%,FAST为31.2%。SALT还开发了动词专用代码,同时保持重建保真度。这些结果表明,机器人动作轨迹提供了语言接地的来源,且在动作表示中保留此结构可显著改善语言条件控制。
英文摘要
Action verbs describe not only the physical outcomes of actions, but also how those actions are performed. Yet action representations in vision-language-action models (VLAs) are typically optimized for reconstruction under L1/L2 losses in raw action space, where numerical proximity need not reflect linguistically meaningful distinctions. On BridgeV2, we show that action trajectories contain verb-grounding information beyond visual state changes, and that reconstruction-only discrete tokenization systematically erodes this information. To address this problem, we introduce SALT, a Semantically ALigned action Tokenizer that augments a VQ-VAE-style tokenizer with an auxiliary objective requiring a frozen vision-language model to recover the episode instruction from quantized action latents. Policies trained with SALT achieve 71.9% average success in SimplerEnv, compared with 42.7% for a reconstruction-only VQ-VAE tokenizer and 31.2% for FAST. SALT also develops verb-specialized codes while maintaining reconstruction fidelity. These results show that robot action trajectories provide a source of language grounding and that preserving this structure in action representations can substantially improve language-conditioned control.
发表机构
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。