arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

V-Link:为视觉-语言-动作模型恢复动作DiT中丢失的视觉表征

V-Link: Recovering Lost Visual Representations in Action DiT for Vision-Language-Action Models

Yehao Lu, Jiarui Yang, Yuning Su, Yufeng Xie, Yu Zhong, Yazhou Zhang, Haiyu Lan, Kaixiang Lu, Peiwen Lin, Chuang Wang, Zequn Qin, Enyu Li, Xi Li

arXiv 2608.25308首次发表:更新:

发表机构

Zhejiang University; AGIBOT; The Hong Kong University of Science and Technology (Guangzhou); Simon Fraser University(浙江大学; AGIBOT; 香港科技大学(广州); 西蒙菲莎大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对VLA模型中动作专家访问视觉信息不足的问题,提出V-Link方法,通过注入空间与语义查询提升动作DiT性能,在多个机器人操作基准及现实任务中均取得显著成功率提升。

AI 中文摘要

视觉-语言-动作(VLA)模型通过整合视觉感知、语言理解和连续动作控制,为通用机器人操作提供了可扩展的路径。然而,我们揭示了VLA架构的一个关键局限:动作专家对VLM特征中可用的3D几何和2D语义信息的访问有限,这种可访问性差距削弱了感知接地并限制了细粒度机器人操作的性能。为解决该问题,我们提出V-Link,其在视觉-语言(VL)到动作(A)的特征转移过程中显式恢复视觉表征。具体而言,V-Link在VLM内部学习互补的空间查询和语义查询表征,并通过非对称路径将其注入动作DiT;语义查询补充原始VLM图像令牌,空间查询为空间接地的动作生成提供专用几何条件。在LIBERO、LIBERO-Plus和RoboTwin 2.0上,我们的V-Link将基础模型GR00T N1.6的平均成功率分别提升了+1.9%、+31.2%和+18.8%;在AGIBOT A3 Ultra上,V-Link在两项现实世界人形机器人任务中进一步分别获得+20%和+24%的提升。

英文摘要

Vision-language-action (VLA) models provide a scalable path toward generalist robotic manipulation by integrating visual perception, language understanding, and continuous action control. However, we reveal a critical limitation of VLA architectures: the action expert has limited access to the 3D geometric and 2D semantic information available in VLM features. This accessibility gap weakens perceptual grounding and limits performance on fine-grained robotic manipulation. To address this issue, we propose V-Link, which explicitly recovers visual representations during the vision-language (VL) to action (A) feature transfer. Specifically, V-Link learns complementary Spatial and Semantic Query representations within the VLM and injects them into Action DiT through asymmetric pathways. Semantic Queries complement the original VLM image tokens, whereas Spatial Queries provide dedicated geometric conditioning for spatially grounded action generation. Across LIBERO, LIBERO-Plus, and RoboTwin 2.0, our V-Link improves the average success rate over base model GR00T N1.6 by +1.9%, +31.2%, and +18.8%, respectively. On the AGIBOT A3 Ultra, V-Link further achieves gains of +20% and +24% on two real-world humanoid tasks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑