发表机构
Noetix Robotics; Tsinghua University(Noetix Robotics; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出DAVIS框架,仅用深度图像和本体感觉实现人形机器人足球接触技能的端到端学习,直接输出25自由度PD目标,并通过仿真和真实实验验证。
AI 中文摘要
人形机器人足球接触技能不仅需要产生高冲击力的脚-球接触:机器人必须在自身运动引起显著视点变化、球频繁从视野中消失以及接触结果不确定的情况下,闭环感知、接近、对齐、冲击和恢复。在这项工作中,我们提出了一个更紧凑且更严格的问题:人形机器人能否仅使用头部安装的深度图像、本体感觉历史以及可选的低维任务命令来学习足球接触技能,并直接输出25自由度关节PD目标,而无需额外的运行时感知或规划模块?为此,我们提出了DAVIS,一个用于人形机器人足球技能的仅深度端到端框架,该框架在训练期间学习可见性感知的辅助几何,并结合GT到预测退火、任务课程和AMP风格的运动先验,以平滑地桥接特权监督和实际部署。基于该框架,我们通过任务特定的对象、命令、奖励和课程定义,实例化了代表性的足球接触技能,包括目标导向射门和定向运球,并通过仿真、Noetix E1真实机器人实验和消融研究进行了验证。
英文摘要
Humanoid soccer contact skills require more than producing high-impact foot-ball contacts: the robot must close the loop over perception, approach, alignment, impact, and recovery while its own motion induces substantial viewpoint changes, frequent loss of the ball from view, and uncertain contact outcomes. In this work, we ask a compact yet stricter question: can a humanoid learn soccer contact skills using only a head-mounted depth image, proprioceptive history, and an optional low-dimensional task command, and directly output 25-DoF joint PD targets without extra runtime perception or planning modules? To this end, we propose DAVIS, a depth-only end-to-end framework for humanoid soccer skills that learns visibility-aware auxiliary geometry during training, and combines GT-to-prediction annealing, task curricula, and AMP-style motion priors to smoothly bridge privileged supervision and real deployment. Built on this framework, we instantiate representative soccer contact skills, including goal-directed shooting and directional dribbling, through task-specific definitions of objects, commands, rewards, and curricula, and validate them through simulation, Noetix E1 real-robot experiments, and ablations.
Comments16 pages, 15 figures. Project page: https://thusi-lab.github.io/DAVIS/