发表机构
Korea University(高丽大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出受被动动态行走启发的框架,通过临时倾斜重力场引导人形机器人学习节能步态,无需参考轨迹,在Unitree G1上降低运输成本6.8-15.2%,并扩展至全向运动。
AI 中文摘要
学习节能的人形机器人运动需要发现机械上经济的步态协调,而不仅仅是减少执行器努力。强化学习通过努力相关的奖励惩罚来促进效率,这仅间接地引导行走的逐步机制。本文提出一个受被动动态行走(PDW)启发的框架,该框架临时创建有利于经济步态发现的等效斜坡条件,并在名义动力学优化之前移除所有PDW特定引导。在早期训练期间,倾斜重力场在平坦碰撞几何体上辅助矢状面前进,并辅以课程耦合的奖励项。核心框架不需要参考轨迹、步态阶段或接触时间表。在29自由度Unitree G1上的五种子向前运动研究中,该框架在指令速度0.5-2.0米/秒范围内将机械运输成本降低6.8-15.2%,且不降低速度跟踪性能。机械功分解将降低归因于正执行器功,奖励匹配比较将引导机制下更快的步态获取与倾斜对收敛经济的额外益处区分开来。该框架扩展到无辅助的全向运动,其中一旦行走特定的运动先验提供运动学协调,其益处持续存在,该组合将速度匹配的运输成本降低18.7%。在硬件上,有运动先验时向前运输成本降低16.3%,无运动先验时降低4.5%,后者在试验间差异范围内。
英文摘要
Learning energy-efficient humanoid locomotion requires discovering mechanically economical gait coordination, not merely reducing actuator effort. Reinforcement learning promotes efficiency through effort-related reward penalties, which guide the step-to-step mechanics of walking only indirectly. This article proposes a framework inspired by passive dynamic walking (PDW) that temporarily creates slope-equivalent conditions favorable to economical gait discovery and removes all PDW-specific guidance before nominal-dynamics optimization. During early training, a tilted-gravity field assists sagittal progression on flat collision geometry, complemented by curriculum-coupled reward terms. The core framework requires no reference trajectories, gait phases, or contact schedules. In a five-seed forward-locomotion study on a 29-DoF Unitree G1, the framework reduces mechanical cost of transport by 6.8-15.2% over commanded speeds of 0.5-2.0m/s without degrading velocity tracking. Mechanical-work decomposition attributes the reduction to positive actuator work, and reward-matched comparisons separate the guided regime's faster gait acquisition from the tilt's additional benefit to converged economy. The framework extends to unassisted omnidirectional locomotion, where its benefit persists once a walking-specific motion prior supplies kinematic coordination, the combination reducing speed-matched cost of transport by 18.7%. On hardware, forward cost of transport falls by 16.3% with the motion prior and by 4.5% without it, the latter within the trial-to-trial spread.