基于步态预测的模型信息安全强化学习用于双足运动
Model-Informed Safe Reinforcement Learning for Bipedal Locomotion via Step-to-Step Prediction
- The Ohio State University(俄亥俄州立大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出基于ALIP模板和DECBF安全证书的模型信息强化学习框架,通过训练整形和运行时动作过滤实现双足运动安全,实验表明减少安全违规但存在安全跟踪权衡。
AI中文摘要:
人形机器人在杂乱、以人为中心的环境中承诺提供多功能移动能力,但实际部署需要原则性的安全保障。经典基于模型的步态生成器能够产生可解释的运动,但往往缺乏现代基于强化学习(RL)方法的鲁棒性和适应性。我们提出了一种以解析角动量线性倒立摆(ALIP)模板为基础的模型信息强化学习框架。我们通过离散指数控制障碍函数(DECBF)为ALIP步进提供逐步安全证书,并将其用作(i)训练时整形信号和(ii)运行时动作过滤器,该过滤器最小程度地调整摆动脚放置以满足模板级约束。全阶安全性在MuJoCo中的Digit人形机器人上使用全身控制器栈进行了实证评估。与无约束基线相比,我们的方法在报告的外部扰动试验中减少了安全违规事件,而较大的横向速度瞬变揭示了安全跟踪的权衡。
英文摘要:
Humanoid robots promise versatile mobility in cluttered, human-centric environments, but real deployment demands principled safety. Classical model-based gait generators yield interpretable motions but often lack the robustness and adaptability of modern reinforcement learning (RL) based approaches. We propose a model-informed reinforcement learning framework anchored to the analytical Angular Momentum Linear Inverted Pendulum (ALIP) template. We provide a step-to-step safety certificate for ALIP stepping via a discrete exponential control barrier function (DECBF) and use it as (i) a training-time shaping signal and (ii) a runtime action filter that minimally adjusts swing-foot placement to satisfy template-level constraints. Full-order safety is evaluated empirically on the Digit humanoid in MuJoCo with a whole-body controller stack. Compared to an unconstrained baseline, our approach reduces safety-violation events in the reported external-disturbance trial, while larger lateral-velocity transients reveal a safety-tracking tradeoff.