arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23483cs.RO

STRIDER:面向人形机器人的基于迈步的多步态分层三维移动操作框架

STRIDER: Stepping-Enabled Multi-Gait Hierarchical 3D Loco-Manipulation Framework for Humanoid Robots

  • X-Humanoid, Humanoid Innovation Department(X-Humanoid,人形机器人创新部门)

机构由 AI 辅助整理,请以论文原文为准。

Yuanzhuo Li, Wen Zhao, Zhe Yong, Xiang Meng, Gang Han, Hengle Ren, Xiaoyang Zheng, Zhen Wang, Yijie Guo

AI总结:

STRIDER提出分层多步态框架,集成三维迈步与全身控制,并引入LD-PPO蒸馏算法,在TianGong Omni上提升落脚点与姿态跟踪精度,实现精确移动操作。

AI中文摘要:

人形机器人的移动操作面临两个显著限制:使用连续速度命令的控制器无法精确调节单个落脚点,而专门的落脚点跟踪模块难以与全身操作集成。此外,标准的基于动作的模仿蒸馏主要转移专家动作,而不明确鼓励异构技能的共享表示。本文介绍了STRIDER,一个分层多步态框架,以弥合这些差距。该框架集成了地形感知的三维迈步逻辑、基于对抗运动先验(AMP)的自然行走以及笛卡尔空间上肢控制:其迈步专家在支撑脚坐标系中选择可行的落脚点,并生成具有间隙感知的摆动轨迹。为了将不同的行走和迈步专家融合为一个可执行的学生策略,我们提出了潜变量蒸馏近端策略优化(LD-PPO),一种通过教师条件潜变量对齐增强的蒸馏算法。通过联合优化策略上的强化学习、基于DAgger的动作重建和潜变量对齐,LD-PPO转移专家动作,同时鼓励跨异构模式的共享技能表示。在TianGong Omni人形机器人上的仿真和真实机器人评估表明,LD-PPO在落脚点跟踪和姿态跟踪精度上优于普通蒸馏PPO。部署在硬件上,STRIDER实现了具有精确落脚点和末端执行器跟踪的多步态移动操作。

英文摘要:

Humanoid loco-manipulation faces two prominent limitations: controllers using continuous velocity commands cannot precisely regulate individual footholds, while specialized foothold-tracking modules are difficult to integrate with whole-body manipulation. Furthermore, standard action-based imitation distillation primarily transfers expert actions, without explicitly encouraging a shared representation of heterogeneous skills. This paper introduces STRIDER, a hierarchical multi-gait framework to bridge these gaps. The framework integrates terrain-aware 3D stepping logic, Adversarial Motion Priors (AMP)-based natural walking, and Cartesian upper-body control: its stepping expert selects feasible footholds in the stance-foot frame and generates clearance-aware swing trajectories. To fuse distinct walking and stepping experts into one executable student policy, we propose Latent Distillation Proximal Policy Optimization (LD-PPO), a distillation algorithm augmented with teacher-conditioned latent alignment. By jointly optimizing on-policy reinforcement learning, DAgger-based action reconstruction, and latent alignment, LD-PPO transfers expert actions while encouraging a shared skill representation across heterogeneous modes. Simulation and real-robot evaluations on the TianGong Omni humanoid show that LD-PPO outperforms vanilla distillation-PPO in foothold-tracking and posture-tracking accuracy. Deployed on hardware, STRIDER realizes multi-gait loco-manipulation with accurate foothold and end-effector tracking.

补充信息

↑