Mind-VLA:面向视觉-语言-动作模型的指令感知空间表示对齐方法
Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models
查看机构详情
- Nanjing University(南京大学)
- Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
- University of Chinese Academy of Sciences(中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文针对现有VLA方法忽略指令指定目标物体3D几何的问题,提出Mind-VLA方法,通过对齐目标物体相关特征实现指令感知3D理解,在LIBERO、CALVIN及真实机器人遮挡任务上性能显著优于基线。
中文摘要 AI 辅助
现有的视觉-语言-动作(VLA)方法通过将其表示与3D场景几何对齐来提升泛化能力,但这些方法本质上是指令无关的:它们均匀对齐整个场景的表示,忽略了语言指令指定的特定目标物体的3D几何,导致在细粒度操作和目标遮挡任务中失效,此类任务的成功依赖于对目标物体而非整个场景的准确3D理解。为解决该问题,本文提出Mind-VLA,一种面向VLA模型的指令感知空间表示对齐方法。具体而言,Mind-VLA首先获取语言指令指定的目标物体,随后准备其目标物体三视图并提取对应的VAE和VGGT特征,最后将VLA模型的潜在表示与这些特征对齐,以实现指令感知的3D理解。Mind-VLA在LIBERO上达到93.9%的性能,在CALVIN上达到4.47的性能,其主干网络仅含3.45亿参数。在含目标遮挡的真实机器人任务中,Mind-VLA的平均成功率达54%,在真实机器人对比中,其性能比表现最佳的指令无关方法高出32个百分点。代码将公开提供。
英文摘要
Recent Vision-Language-Action (VLA) methods improve generalization by aligning their representations with 3D scene geometry. However, these methods are fundamentally instruction-agnostic: the representations align the entire scene uniformly, neglecting the 3D geometry of the specific target object designated by the language instruction. This causes failures on fine-grained manipulation and target occlusion tasks, where success depends on accurate 3D understanding of the target object rather than the entire scene. To address this, we present Mind-VLA, an instruction-aware spatial representation alignment method for VLA models. Specifically, Mind-VLA first obtains the target object specified by the language instruction, then prepares its canonical target views and extracts the corresponding VAE and VGGT features. Finally, the latent representation of the VLA model is aligned with these features to enable instruction-aware 3D understanding. Mind-VLA reaches 94.4% on LIBERO and 4.47 on CALVIN with a compact 345M-parameter backbone. On real-robot tasks with target occlusion, Mind-VLA reaches 54% average success, outperforming the matched scene-VGGT control by 26 percentage points.