欠驱动双足机器人无碰撞运动的分层强化学习
Hierarchical Reinforcement Learning for Collision-Free Locomotion of an Underactuated Biped
浏览论文内容
中文总结 AI 辅助
针对欠驱动双足机器人的避障运动,提出分层强化学习框架,高层输出速度指令、低层跟踪执行,联合SAC训练,在随机环境中成功率显著优于规划器混合基线。
中文摘要 AI 辅助
双足机器人无法在不干扰其平衡的情况下偏离路径以避开障碍物,而这种耦合在欠驱动平台上最为严重,例如本文考虑的双足机器人,其每条腿有四个驱动关节,且没有髋部或踝部的滚动关节。本文提出了一种分层强化学习(HRL)框架,其中高层(HL)策略观察机器人姿态、36个射线投射邻近测量、移动障碍物状态以及一个滚动时域局部目标,并每十个控制步输出一个身体速度指令$(v_x, v_y, \omega_{yaw})$,而速度条件化的低层(LL)策略通过PD控制的关节目标跟踪每个指令。两个策略通过软演员-评论家(SAC)算法联合训练。由于收敛的步态是任务无关的,它被冻结并通过经典规划器在同一指令接口上驱动,产生三个受控基线:SAC+A*、SAC+RRT*和SAC+APF。在随机化的PyBullet环境中,每种方法进行100次评估试验,所提出的方法在静态试验中达到目标的成功率为98.0%,在动态试验中为88.0%,而规划器混合方法最多分别为78.0%和68.0%,路径长度在A*参考值的4%以内,消融实验证实每个观察通道和奖励项都对这一性能有实质性贡献。
英文摘要
A bipedal robot cannot deviate from its path to avoid an obstacle without disturbing its balance, and this coupling is most severe on underactuated platforms such as the biped considered here, which has four actuated joints per leg and no hip or ankle roll. This paper presents a Hierarchical Reinforcement Learning (HRL) framework in which a High-Level (HL) policy observes the robot pose, 36 raycast proximity measurements, moving-obstacle states, and a receding-horizon local goal, and outputs a body-velocity command $(v_x, v_y, ω_{yaw})$ every ten control steps, while a velocity-conditioned Low-Level (LL) policy tracks each command through PD-controlled joint targets. Both policies are trained jointly with Soft Actor-Critic (SAC). Because the converged gait is task-agnostic, it is frozen and driven by classical planners over the same command interface, yielding three controlled baselines: SAC+A*, SAC+RRT*, and SAC+APF. Across 100 evaluation trials per method in randomized PyBullet environments, the proposed method reaches the goal in 98.0% of static and 88.0% of dynamic trials, against at most 78.0% and 68.0% for the planner hybrids, with path lengths within 4% of the A* reference, and ablations confirm that each observation channel and reward term contributes materially to this performance.
发表机构
- Indian Institute of Technology Mandi(印度理工学院曼迪分校)
机构由 AI 辅助整理,请以论文原文为准。