UMI-Bridge:人类与机器人操作数据中的动作锚定潜在对齐
UMI-Bridge: Action-Anchored Latent Alignment across Human and Robot Manipulation Data
浏览论文内容
中文总结 AI 辅助
UMI-Bridge利用UMI作为中间域,通过动作锚定潜在对齐,在无机器人演示下训练双视角潜在动作模型,实现数据高效的机器人学习,平均成功率提升至91.7%。
中文摘要 AI 辅助
真实机器人演示数据有限,这促使研究者利用无需机器人即可收集的人类操作数据,包括第一人称视频和手持式通用操作接口(UMI)演示。然而,视角、具身形态以及可用动作监督的差异,使得难以根据操作动作而非视觉外观来对齐这些数据源之间的表征。我们提出UMI-Bridge,该方法利用UMI作为中间域,根据动作等价性而非像素相似性来对齐表征。UMI动作监督将潜在表征锚定到末端执行器运动和夹爪行为,同时同步的头戴-手腕观测和配对的自我中心-UMI片段支持跨视角和跨域对齐。我们在无需机器人演示的人类操作数据上训练了一个双视角潜在动作模型(LAM),然后冻结其手腕教师模型和动力学模型,以规范在UMI和机器人数据上的视觉-语言-动作(VLA)后训练。共享的手腕接口使得这种训练时监督能够跨越两个域,同时保留策略的标准推理架构。在三个真实机器人任务中,UMI-Bridge的平均成功率达到91.7%,而使用匹配的UMI和机器人数据的朴素协同训练为73.3%。在两个数据效率任务中,它仅使用25%的机器人演示结合UMI数据,就超越了使用全部数据的仅机器人基线。此外,在另外两个仅从UMI演示学习且无任务特定机器人演示的任务中,它分别达到了85%和90%的成功率。这些结果支持动作锚定的潜在对齐用于数据高效的机器人学习和UMI到机器人的任务迁移。
英文摘要
Real-robot demonstrations are limited, motivating the use of human manipulation data collected without robots, including egocentric videos and handheld Universal Manipulation Interface (UMI) demonstrations. However, differences in viewpoint, embodiment, and available action supervision make it difficult to align representations across these sources according to manipulation motion rather than visual appearance. We introduce UMI-Bridge, which uses UMI as an intermediate domain to align representations according to action equivalence rather than pixel similarity. UMI action supervision anchors the latent representation to end-effector motion and gripper behavior, while synchronized head-wrist observations and paired ego-UMI clips support alignment across views and domains. We train a dual-view latent action model (LAM) on human manipulation data without robot demonstrations, then freeze its wrist teacher and dynamics model to regularize vision-language-action (VLA) post-training on UMI and robot data. The shared wrist interface enables this training-time supervision across both domains while preserving the policy's standard inference architecture. Across three real-robot tasks, UMI-Bridge achieves 91.7% mean success versus 73.3% for Naive Co-training with matched UMI and robot data. On two data-efficiency tasks, it surpasses a full-data Robot-only baseline using 25% of the robot demonstrations together with UMI data. It also achieves 85% and 90% success on two additional tasks learned from UMI demonstrations without task-specific robot demonstrations. These results support action-anchored latent alignment for data-efficient robot learning and UMI-to-robot task transfer.
发表机构
- Tsinghua University(清华大学)
- Simple AI
机构由 AI 辅助整理,请以论文原文为准。