arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07619cs.RO

GWM-VLA:面向视觉-语言-动作学习的几何感知隐式世界建模

GWM-VLA: Geometry-Aware Latent World Modeling for Vision-Language-Action Learning

Yanping Zhao, Hang Yu, Yiwei Wang, Chen Ye, Siyu Tian, Di Zhang, Qingjun Wang, Qian Chen, Junqiao Zhao, Chen Ye, Guang Chen

首次发表
浏览论文内容

中文总结 AI 辅助

针对VLA模型在视觉/环境变化下性能下降问题,提出GWM-VLA框架,通过几何感知多视角编码等技术提升模型鲁棒性,经仿真和真实环境实验验证有效。

中文摘要 AI 辅助

视觉-语言-动作(VLA)模型在机器人操纵任务中表现出色,但在视觉和环境发生变化时性能往往会下降。隐式世界建模为提升模型鲁棒性提供了一种有前景的方法,不过现有方法通常独立编码相机视角,且在预测整体场景动态时未显式建模其几何关系。我们提出GWM-VLA,一种面向VLA学习的几何感知隐式世界建模框架。GWM-VLA结合了几何感知多视角状态编码、全局上下文条件目标视角预测,以及由机器人动作监督所确立的共享隐式-动作表示。具体而言,VGGT-Ω在每个时间步联合聚合多视角观测结果,以构建几何感知多视角状态。该隐式世界模型利用多视角聚合后得到的patch和register token,预测所选目标视角的下一步patch token,从而在不预测完整多视角状态的情况下保留多视角几何信息。我们在实验中采用手腕视角作为目标,更侧重于末端执行器运动和末端执行器-物体的局部交互。最后,共享隐式-动作表示同时对隐式世界模型和流匹配动作头进行条件约束,使得隐式预测监督和真实机器人动作监督能够共同塑造同一隐式-动作表示。在仿真和真实环境下开展的实验证明了GWM-VLA的有效性和鲁棒性。

英文摘要

Vision-Language-Action (VLA) models achieve strong robotic manipulation performance but often degrade under visual and environmental shifts. Latent world modeling offers a promising approach to improving robustness, yet existing methods commonly encode camera views independently and predict holistic scene dynamics without explicitly modeling their geometric relationships. We propose GWM-VLA, a geometry-aware latent world modeling framework for VLA learning. GWM-VLA combines geometry-aware multi-view state encoding, global context-conditioned target-view prediction, and shared latent-action representations grounded by robot-action supervision. Specifically, VGGT-$Ω$ jointly aggregates multi-view observations at each timestep to construct geometry-aware multi-view states. The latent world model predicts the next-step patch tokens of a selected target view using patch and register tokens obtained after multi-view aggregation, thereby retaining multi-view geometric information without predicting the complete multi-view state. We use the wrist view as the target in our experiments, placing greater emphasis on end-effector motion and local gripper-object interactions. Finally, the shared latent-action representations condition both the latent world model and the flow-matching action head, allowing latent-prediction supervision and ground-truth robot-action supervision to jointly shape the same latent-action representations. Experiments across both simulation and real-world environments demonstrate the effectiveness and robustness of GWM-VLA.

↑