arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GeoLAM:从无标注人类视频中学习几何基础的潜在动作

GeoLAM: Learning Geometry-Grounded Latent Actions from Unlabeled Human Videos

Yifan Xie, Hekun Tian, Jinkun Liu, YuAn Wang, Qiao Sun, Wenbo Ding

arXiv 2609.17099首次发表:更新:

发表机构

Tsinghua University; Beijing Institute of Technology; Fudan University(清华大学; 北京理工大学; 复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

GeoLAM从无标注人类视频中学习几何基础的潜在动作,结合几何特征与4D教师监督,提升机器人操作性能。

AI 中文摘要

人类视频提供了丰富的操作经验,但提取保留有用运动的动作表征仍然具有挑战性。仅靠视觉重建可能会将操作相关的运动与外观变化和相机移动纠缠在一起。我们提出了GeoLAM,一个从无动作人类视频中学习几何基础潜在动作的框架。GeoLAM通过冻结的几何特征层级进行未来帧重建,并结合来自仅用于训练的4D几何教师模型的运动监督。几何表征提供了结构先验,而教师模型的预测产生了捕获3D位移、残差图像平面运动和表面方向变化的空间池化目标。可见性和置信度加权减少了不可靠估计的贡献,鼓励连续的潜在动作保留几何运动,而无需显式的手部姿态或手部轨迹标注。在无动作标签的视频预训练之后,学习到的表征为在世界动作模型上训练提供了转移目标,该模型在带有动作标签的机器人演示上进行训练。该模型联合去噪潜在动作和可执行的动作块,未来视频预测仅用作辅助训练任务。因此,部署既不需要几何教师模型,也不需要未来视频生成。在潜在动作基准和机器人操作任务上的评估证明了GeoLAM的强劲性能。

英文摘要

Human videos provide rich manipulation experience, but extracting action representations that preserve useful motion remains challenging. Visual reconstruction alone can entangle manipulation-related motion with appearance changes and camera movement. We present GeoLAM, a framework for learning geometry-grounded latent actions from action-free human videos. GeoLAM combines future-frame reconstruction through a frozen geometric feature hierarchy with motion supervision from a training-only 4D geometry teacher. The geometric representation provides a structural prior, while the teacher's predictions yield spatially pooled targets capturing 3D displacement, residual image-plane motion, and surface-orientation changes. Visibility and confidence weighting reduces the contribution of unreliable estimates, encouraging continuous latent actions to retain geometric motion without explicit hand-pose or hand-trajectory annotations. After video pretraining without action labels, the learned representation provides transition targets for a world-action model trained on action-labeled robot demonstrations. The model jointly denoises latent actions and executable action chunks, with future-video prediction used only as an auxiliary training task. Deployment therefore requires neither the geometry teacher nor future-video generation. Evaluations on a latent-action benchmark and robotic manipulation tasks demonstrate the strong performance of GeoLAM.

Comments8 pages, 6 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑