arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ViLoMan:学习人形机器人视觉-本体感觉全身移动操作技能

ViLoMan: Learning Visual-Proprioceptive Whole-Body Loco-Manipulation Skills for Humanoid Robots

Zejie Tian, Ruibing Hou, Bingpeng Ma, Börje F. Karlsson, Shiguang Shan

arXiv 2609.19340首次发表:更新:

发表机构

State Key Laboratory of AI Safety, Institute of Computing Technology, CAS; University of Chinese Academy of Sciences (CAS); Beijing Academy of Artificial Intelligence (BAAI)(中国科学院计算技术研究所人工智能安全国家重点实验室; 中国科学院大学; 北京智源人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ViLoMan提出可扩展框架,通过将人-物交互演示转为机器人轨迹并利用师生蒸馏学习统一策略,使Unitree G1人形机器人仅凭机载感知完成关门任务,实现稳健泛化与仿真到现实迁移。

AI 中文摘要

人形移动操作需要自适应全身协调,以无缝集成运动与物理交互。尽管近期取得了进展,但由于缺乏多样化、物理可执行的机器人-物体交互数据,以及难以直接从机载观测中学习统一的全身控制,自主移动操作的学习仍然具有挑战性。我们提出了ViLoMan,一个用于自主人形移动操作的可扩展框架。ViLoMan首先将人-物交互的部分运动学演示转换为完整、物理可执行的机器人轨迹。然后,它利用这些轨迹在教师-学生蒸馏框架中学习一个统一策略,该策略直接将深度视觉观测和本体感觉测量映射到关节级全身动作。在部署时,该策略既不需要参考运动,也不需要中间命令。我们在仿真和现实世界中,针对多样化的门配置和机器人初始条件,对ViLoMan进行了关门任务评估。实验结果表明,单一策略使Unitree G1人形机器人仅使用机载深度感知和本体感觉即可完成整个任务,同时在不同任务变化中稳健泛化,并有效实现从仿真到现实的迁移。项目页面:此HTTP URL。

英文摘要

Humanoid loco-manipulation requires adaptive whole-body coordination to seamlessly integrate locomotion and physical interaction. Despite recent advances, learning autonomous loco-manipulation remains challenging due to the scarcity of diverse, physically executable robot-object interaction data and the difficulty of learning unified whole-body control directly from onboard observations. We present ViLoMan, a scalable framework for autonomous humanoid loco-manipulation. ViLoMan first transforms partial kinematic demonstrations of human-object interactions into complete, physically executable robot trajectories. It then leverages these trajectories within a teacher-student distillation framework to learn a unified policy that maps egocentric depth observations and proprioceptive measurements directly to joint-level whole-body actions. During deployment, the policy requires neither reference motions nor intermediate commands. We evaluate ViLoMan on door-closing tasks across diverse door configurations and robot initial conditions in both simulation and the real world. Experimental results demonstrate that a single policy enables a Unitree G1 humanoid to complete the full task using only onboard depth sensing and proprioception, while generalizing robustly across task variations and transferring effectively from simulation to reality. Project page: viloman-anonymous.pages.dev.

Comments10 pages, 7 figures. Project page: https://viloman-anonymous.pages.dev/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑