arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21948cs.ROcs.CV

GALA:面向跨本体视觉-语言-动作模型预训练的几何感知潜在动作建模

GALA: Geometry-Aware Latent Action Modeling for Vision-Language-Action Model Pretraining across Embodiments

发表机构清华大学交叉信息研究院 · 上海期智研究院
查看机构详情
  • Institute for Interdisciplinary Information Sciences, Tsinghua University(清华大学交叉信息研究院)
  • Shanghai Qi Zhi Institute(上海期智研究院)

机构由 AI 辅助整理,请以论文原文为准。

Yichen Liu, Puzhen Yuan, Xiang Zhu, Yanjiang Guo, Jianyu Chen

首次发表
浏览论文内容

中文总结 AI 辅助

GALA提出几何感知潜在动作建模框架,通过统一末端执行器运动表示增强跨本体VLA预训练,有效捕捉细粒度关节运动,在RoboCasa-GR1和真实世界任务中分别取得68.3%和75.5%的成功率。

中文摘要 AI 辅助

从多本体数据集中学习大规模视觉-语言-动作(VLA)模型,由于不同末端执行器之间的动作空间异构,仍然具有挑战性。尽管潜在动作模型(LAMs)能够从多样的视频数据中学习与本体无关的动作表示,但现有的基于图像的LAMs往往难以捕捉细粒度的末端执行器关节运动,特别是人类和灵巧机器人手部的手指级几何变化。为解决这一局限,我们提出了GALA,一种几何感知潜在动作建模框架,该框架将3D末端执行器几何运动增强到基于图像的潜在动作中。然而,简单地将点云纳入会生成共享语义有限的细粒度动作表示,从而阻碍跨本体预训练。为解决此问题,我们引入了统一末端执行器运动表示(UEMR),该表示在保留细粒度运动信息的同时,提高了潜在动作的跨本体泛化能力。基于UEMR,GALA将捕捉场景级动态的视觉潜在动作与捕捉共享细粒度末端执行器关节运动的几何潜在动作相结合,为从多本体数据(包括无动作的以自我为中心的人类视频)进行VLA预训练提供了有效的监督。在细粒度运动探测、跨本体检索和下游VLA评估上的实验证明了GALA在跨本体建模可泛化的细粒度运动方面的有效性,在RoboCasa-GR1上达到了68.3%的成功率,在真实世界中达到了75.5%的成功率。代码、附录和演示可在该https URL获取。

英文摘要

Learning large-scale vision-language-action (VLA) models from multi-embodiment datasets remains challenging due to heterogeneous action spaces across end effectors. Although latent action models (LAMs) can learn embodiment-agnostic action representations from diverse video data, existing image-based LAMs often fail to capture fine-grained end-effector articulation, particularly finger-level geometric changes in human and dexterous robot hands. To address this limitation, we propose GALA, a Geometry-Aware Latent-Action modeling framework that augments image-based latent actions with 3D end-effector geometric motion. However, naively incorporating point clouds yields fine-grained action representations with limited shared semantics, hindering cross-embodiment pretraining. To address this issue, we introduce the Unified End-effector Motion Representation (UEMR), which preserves fine-grained motion information while improving the cross-embodiment generalizability of latent actions. Building upon UEMR, GALA combines visual latent actions that capture scene-level dynamics with geometric latent actions that capture shared fine-grained end-effector articulation, providing effective supervision for VLA pretraining from multi-embodiment data, including action-free ego-centric human videos. Experiments on fine-grained motion probing, cross-embodiment retrieval, and downstream VLA evaluation demonstrate GALA's effectiveness in modeling generalizable fine-grained motions across embodiments, achieving 68.3% RoboCasa-GR1 success rate and 75.5% real-world success rate. Code, appendix, and demos are available at https://puzhenyuan.github.io/GALA-website/.

↑