arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

迈向通用人形移动操作模型:基于以自我为中心全身人类数据预训练

Towards a General Humanoid Loco-Manipulation Model via Egocentric Whole-Body Human Data Pretraining

Chongyang Xu, Zhao Wu, Jin Chen, Yiming Jiang, Jinhui Ye, Yuming Jiang, Shifeng Zhang, Ziliang Feng, Mu Xu, Yilun Chen, Li Lu, Steven C. H. Hoi

arXiv 2610.00438首次发表:更新:

发表机构

Alibaba Group; Sichuan University; Shanghai Innovation Institute; Beihang University; Hong Kong University of Science and Technology(阿里巴巴集团; 四川大学; 上海创新研究院; 北京航空航天大学; 香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出利用以自我为中心的人类全身数据预训练,构建通用人形移动操作模型λ₀,通过三阶段训练实现高效迁移,并在多个任务上取得最先进性能。

AI 中文摘要

人形全身操作技术已取得快速进展,使策略能够协调运动、姿态、双臂交互和灵巧手部动作。与此同时,以自我为中心的人类视频提供了跨物体和场景的多样化日常交互示例,无需机器人操作即可提供可扩展的监督信号。然而,这些视频现有的监督信号对全身运动以及与手-物交互的协调覆盖有限,而通过人形遥操作获取此类监督信号既昂贵又难以扩展。因此,我们探索人类经验如何支持人形移动操作的可扩展学习。为支持这项研究,我们引入了HumanVerse-500,这是一个包含500小时多样化人类移动操作行为的数据集,这些行为在开放世界环境中采集,使用轻量级可穿戴系统同步以自我为中心的视频与身体和手部运动。基于该数据集,我们开发了λ₀,一个全身人形视觉-语言-动作策略,通过三阶段训练实现:首先从多样化以自我为中心的数据集中学习交互,然后利用HumanVerse-500协调身体和手部运动,最后将策略适应下游任务和机器人实体。在这些阶段中,λ₀学习一个共享表示空间用于人类经验迁移,而特定领域的接口处理人类与机器人状态和动作之间的差异。我们在SIMPLE和4个真实世界移动操作任务上评估λ₀,取得了最先进的性能,并进一步分析其扩展行为、泛化能力和训练阶段贡献,以理解人类数据如何支持下游全身人形控制。我们将发布代码、模型和数据以支持进一步研究。

英文摘要

Humanoid whole-body manipulation has advanced rapidly, enabling policies to coordinate locomotion, posture, bimanual interaction, and dexterous hand movements. Meanwhile, egocentric human videos provide diverse examples of everyday interactions across objects and scenes, offering scalable supervision without robot operation. However, existing supervision from these videos provides limited coverage of whole-body movement and coordination with hand-object interaction, while obtaining such supervision through humanoid teleoperation is also costly and difficult to scale. We therefore explore how human experience can support scalable learning of humanoid loco-manipulation. To support this study, we introduce HumanVerse-500, a 500-hour dataset of diverse human loco-manipulation behaviors in open-world environments, collected with a lightweight wearable system that synchronizes egocentric video with body and hand motion. Building on this dataset, we develop $λ_0$, a whole-body humanoid vision-language-action policy, through three-stage training that first learns interaction from diverse egocentric datasets, then coordinates body and hand motion using HumanVerse-500, and finally adapts the policy to downstream tasks and robot embodiments. Across these stages, $λ_0$ learns a shared representation space for human experience transfer, while domain-specific interfaces handle differences between human and robot states and actions. We evaluate $λ_0$ on SIMPLE and 4 real-world loco-manipulation tasks, achieving state-of-the-art performance, and further analyze its scaling behavior, generalization, and training-stage contributions to understand how human data support downstream whole-body humanoid control. We will release our code, models, and data to support further research.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑