发表机构
California Institute of Technology(加州理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出两层人形机器人运动架构,结合流匹配生成器与感知跟踪策略,通过离线强化学习微调提升地形穿越和技能选择性能。
AI 中文摘要
通用人形机器人需要具备多技能、感知能力、动态性和鲁棒性的 locomotion 控制器,以便能够到达人类所能到达的任何地方。在这项工作中,我们提出了一种两层 locomotion 架构:(1)一个感知型流匹配运动生成器,从原始深度图像中规划全身轨迹;(2)一个使用控制引导强化学习训练的感知型跟踪策略,跟踪这些运动。两个策略都在一个由动态优化的人类数据创建的地形一致运动片段库上进行训练,这产生了准确的 velocity 跟踪和地形一致的参考。我们的核心贡献是一个简单但有效的 off-policy 强化学习微调循环,用于改进运动生成器。使用结构化搜索方法与生成器一起收集数据,用于优势加权回归。这个 off-policy 循环比 on-policy 残差微调更具样本效率,并提高了在未见过的几何形状和技能组合上的地形一致性。我们发现,成功的地形穿越率提高了多达 25 个百分点,技能选择提高了多达 80 个百分点。通过使用原始深度图像感知环境,不需要里程计或高度图,并且易于户外部署。使用两个摄像头,策略可以更远地看到地形,并根据命令速度调整其速度,从而穿越地形。一个策略对使 Unitree G1 人形机器人能够在户外环境中行走、跑步、站立、跳上和跳下箱子以及穿越楼梯。项目页面:此 https URL
英文摘要
General purpose humanoids require locomotion controllers that are multi-skill, perceptive, dynamic, and robust enough to go anywhere humans can. In this work, we present a two layer locomotion architecture: (1) a perceptive flow matching motion generator plans whole body trajectories from raw depth images while a (2) perceptive tracking policy trained with control-guided RL follows these motions. Both policies are trained on a library of terrain consistent motion clips created with dynamically optimized human data which yields both accurate velocity tracking and terrain consistent references. Our central contribution is a simple yet effective off-policy RL fine tuning loop that improves the motion generator. A structured search method is used with the generator to gather data for advantage weighted regression. This off-policy loop is much more sample efficient than on-policy residual fine tuning and improves terrain consistency on unseen geometries and skill compositions. We find that successful terrain traversals increased by up to 25 percentage points and skill selection improved by up to 80 percentage points. By using raw depth images to perceive the environment no odometry or height maps are needed, and outdoor deployment is easy. With two cameras, the policy can see terrain coming from further away and adjust its velocity regardless of the commanded speed so it can traverse the terrain. A single policy pair enables a Unitree G1 humanoid to walk, run, stand, jump on and off of boxes, and traverse stairs in outdoor environments. Project page: https://zolkin1.github.io/generate-track-improve/
Comments12 pages