arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03388cs.RO

功夫运动员机器人:从视频学习高动态人形运动并具备统一鲁棒恢复能力

KungfuAthleteBot: learning high-dynamic humanoid motion from video with unified robust recovery

Zhongxiang Lei, Lulu Cao, Xuyang Wang, Tianyi Qian, Jinyan Liu, Xuesong Li

首次发表
浏览论文内容

中文总结 AI 辅助

提出KungfuAthleteBot框架,通过物理校正、伪低动能采样和统一训练策略,从视频学习高动态人形运动,实现约0.7秒的快速跌倒恢复。

中文摘要 AI 辅助

视频是丰富且廉价的人体运动数据来源,其中包含大量极限运动行为。然而,使其可用于人形机器人并非简单地将重建轨迹进行重定向:视频衍生的运动在物理上不一致,缺乏驱动信息,并且对失败或恢复毫无涉及。我们提出功夫运动员机器人(KungfuAthleteBot, KAB),一个将学习高动态运动作为核心问题并依次解决上述三个失败模式的框架。(C1)我们构建了功夫运动员数据集,该数据集来源于国家级武术运动员的视频,并引入物理引导的抛物线轨迹校正,以消除重建的空中和落地阶段中的高度漂浮、地面穿透和高频抖动。(C2)由于视频不包含力信息,对重建轨迹的严格跟踪在动态上是不可行的,而基于误差的初始化会不断从不可行的空中姿态重新启动策略。我们引入物理驱动的伪低动能(LKE)采样,作为使此类参考可学习的关键机制:它将初始化偏向动态可行的状态,使策略能够发现可行的驱动模式,而不是模仿不可行的模式。(C3)最后,我们引入一种直接训练范式,其中干扰抑制和跌倒恢复在与跟踪视频运动的同一策略内学习,无需恢复参考数据,也无需手动模式切换。在人形机器人上,KAB能从视频中学习动态技能,并在约0.7秒内从任意跌倒中恢复,这是统一策略中报告的最快恢复时间。对统一策略的消融实验证实了其组件的必要性,支持了以下观点:修复和补偿视频数据,而不仅仅是收集更多数据,才是解锁高动态人形技能的关键。

英文摘要

Video is an abundant, inexpensive source of human motion data that is rich in extreme athletic behaviors. Making it usable for humanoid robots, however, is not a matter of simply retargeting a reconstructed trajectory: video-derived motion is physically inconsistent, devoid of actuation information, and says nothing about failure or recovery. We present KungfuAthleteBot (KAB), a framework that treats learning high-dynamic motion from video as the central problem and resolves each of these three failure modes in turn. (C1) We build the KungfuAthlete dataset from videos of national-level martial artists and introduce a physics-guided parabolic trajectory correction that removes height floating, ground penetration, and high-frequency jitter from reconstructed aerial and landing phases. (C2) Because video carries no force information, strict tracking of a reconstructed trajectory is dynamically infeasible, and error-driven initialization keeps re-launching the policy from infeasible aerial poses. We introduce physics-driven pseudo-low-kinetic-energy (LKE) sampling, our central mechanism for making such references learnable: it biases initialization towards dynamically feasible states, letting the policy discover feasible actuation patterns instead of imitating infeasible ones. (C3) Finally, we introduce a direct training paradigm in which disturbance rejection and fall recovery are learned inside the same policy that tracks the video motion, requiring no recovery reference data and no manual mode switching. On a humanoid robot, KAB learns dynamic skills from video and recovers from arbitrary falls in about 0.7 s, the fastest reported recovery for a unified policy. Ablations on the unified policy confirm the necessity of its components, supporting the view that repairing and compensating video data, rather than only collecting more of it, is what unlocks high-dynamic humanoid skills.

发表机构

  • Beijing Institute of Technology(北京理工大学)
  • QIYUAN Lab(启元实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑