arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向多样地形的四足机器人分层强化学习节能步态自适应

Energy-Efficient Gait Adaptation via Hierarchical Reinforcement Learning for Quadrupedal Locomotion Across Diverse Terrains

Ammar Issa, Anubhav Singh, Anton Tsaritsin, Sergey Kolyubin

arXiv 2610.10297首次发表:更新:

发表机构

ITMO University; Biomechatronics and Energy-Efficient Robotics Lab (BE2R)(ITMO大学; 仿生与节能机器人实验室(BE2R))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出一种分层强化学习框架,分离高频运动执行与低频步态优化,在多种地形和速度下降低运输成本,并实现零样本仿真到现实迁移。

AI 中文摘要

虽然能量效率是腿式机器人运动控制的一个关键目标,但在不同速度范围和地形条件下实现低能耗同时保持稳健性能仍然是一个关键挑战。这对于端到端强化学习策略尤其如此,其中步态生成、运动执行和能量优化紧密耦合,导致对奖励设计高度敏感。在这项工作中,我们提出了一个分层强化学习(HRL)框架,将用于稳定且稳健的关节级运动执行的高频策略与显式最小化运输成本(CoT)的低频步态自适应分离。基于Isaac的三阶段训练流程实现了零样本仿真到现实迁移,并提高了跟踪精度、稳健性和能量效率。学习到的层级结构表现出自动的与速度相关的步态自适应,从低速时的踱步步态过渡到高速时的疾驰步态。我们在仿真中针对代表性的单策略和分层运动基线验证了所提出的方法,在广泛的指令速度范围内展示了降低的CoT,同时在平坦、不平坦粗糙和倾斜地形上保持了稳健的运动。我们进一步通过在物理Unitree AlienGo四足机器人上的零样本部署证明了其实用可行性。

英文摘要

While energy efficiency is a critical objective for legged-robot locomotion control, achieving low energy consumption while maintaining robust performance across different velocity ranges and terrain conditions remains a key challenge. This is particularly true for end-to-end RL policies, where gait generation, motion execution, and energy optimization are tightly coupled, leading to high sensitivity to reward design. In this work, we propose a hierarchical reinforcement learning (HRL) framework that separates a high-frequency policy for stable and robust joint-level motion execution from low-frequency gait adaptation that explicitly minimizes the cost of transport (CoT). The three-stage Isaac-based training procedure enables zero-shot sim-to-real transfer with improved tracking accuracy, robustness, and energy efficiency. The learned hierarchy exhibits automatic speed-dependent gait adaptation, transitioning from pacing at low speeds to trotting at higher speeds. We validate the proposed approach in simulation against representative single-policy and hierarchical locomotion baselines, demonstrating reduced CoT over a broad range of commanded velocities, while maintaining robust locomotion across flat, uneven rough, and inclined terrains. We further demonstrate its practical feasibility through zero-shot deployment on a physical Unitree AlienGo quadruped.

Comments9 pages. Submitted to IEEE ICRA 2027. Ammar Issa, Anubhav Singh, and Anton Tsaritsin contributed equally

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑