arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29850cs.ROcs.CV

BeyondRetarget:直接从单目视频学习可执行的人形机器人动作

BeyondRetarget: Learning Executable Humanoid Motions Directly from Monocular Video

  • Nanjing University(南京大学)
  • Jiangsu Mobile Information System Integration Co., Ltd.(江苏移动信息系统集成有限公司)
  • China Mobile Zijin (Jiangsu) Innovation Research Institute Co., Ltd.(中国移动紫金(江苏)创新研究院有限公司)

机构由 AI 辅助整理,请以论文原文为准。

Tianyu Xiong, Yi Lu, Jinrui Wang, Ziqi Liang, Dandan Lei, Xiaoyang Zhou, Xiao-xiao Long, Qiu Shen, Xun Cao

AI总结:

BeyondRetarget提出端到端框架,直接从单目视频学习机器人动作,摒弃人类表示,结合接触感知优化,提升执行成功率与鲁棒性。

AI中文摘要:

从人类视频中学习可执行的动作,为人类形机器人获取示范动作提供了一种可扩展的解决方案。然而,现有的流程通常首先构建一个显式的人类运动表示,然后通过运动重定向将其转换为机器人动作。尽管此类方法可以有效利用大量现有的人类数据进行训练,但人类与类人机器人在运动机制和关节自由度配置上的显著差异,使得这种以人类表示为中心的方法生成的动作难以在机器人上执行。此外,人类运动估计过程中引入的误差不可避免地传播到重定向阶段,且无法通过联合优化来消除。我们提出了BeyondRetarget,一个端到端的框架,直接将单目RGB视频映射到机器人动作。该框架摒弃了显式的人类表示,直接从视觉观测中学习面向机器人的隐式表示,使模型能够捕获跨形态的运动结构。为了生成更适合机器人执行的动作,我们进一步设计了一种接触感知的运动优化机制,以提高时间一致性和物理合理性。实验表明,BeyondRetarget显著提高了生成机器人动作的准确性和鲁棒性,同时在仿真环境和真实类人机器人上实现了更高的执行成功率和更低的延迟。

英文摘要:

Learning executable motions from human videos offers a scalable solution for humanoid robots to acquire demonstration motions. However, existing pipelines typically first construct an explicit human motion representation and then convert it into robot motions via motion retargeting. Although such methods can effectively leverage large volumes of existing human data for training, the substantial differences between humans and humanoid robots in locomotion mechanisms and joint degree-of-freedom configurations make motions generated by this human-representation-centric approach difficult to execute on robots. Furthermore, errors introduced during human motion estimation inevitably propagate to the retargeting stage and cannot be eliminated via joint optimization. We propose BeyondRetarget, an end-to-end framework that directly maps monocular RGB videos to robot motions. Discarding the explicit human representation, this framework learns robot-oriented implicit representations directly from visual observations, enabling the model to capture cross-morphology motion structures. To generate motions more suitable for robot execution, we further design a contact-aware motion optimization mechanism to improve temporal consistency and physical plausibility. Experiments show that BeyondRetarget significantly improves the accuracy and robustness of generated robot motions, while achieving higher execution success rates and lower latency in both simulation environments and real humanoid robots.

↑