基于时序逻辑规范的学习步态感知四足运动
Learning Gait-Aware Quadruped Locomotion with Temporal Logic Specifications
浏览论文内容
中文总结 AI 辅助
提出用信号时序逻辑参数化约束指定步态,通过奖励塑形提供密集连续奖励,在Barkour机器人上实现更紧的速度跟踪和稳定训练。
中文摘要 AI 辅助
四足运动的强化学习通常依赖于固定的、手工设计的马尔可夫奖励函数,这限制了学习策略的可解释性,并且缺乏对步态行为的显式控制。我们引入了一个框架,其中不同的步态使用信号时序逻辑(STL)中表达的参数化约束来指定。这些包括安全界限、步态同步约束、命令跟踪和驱动界限。根据这些规范,我们开发了一种奖励塑形机制,为学习代理提供密集、连续的奖励景观,编码期望的行为。我们为三种速度模式(走-小跑、小跑、跳跃)定义了参数化STL模板,从参考轨迹中校准其参数,并使用STL鲁棒性的平滑近似在轨迹上计算奖励。生成的奖励可用于提供与近端策略优化(PPO)兼容的塑形梯度。我们在Google的Barkour四足机器人上实例化了该方法,使用MuJoCo XLA(MJX)。我们利用模拟器内的并行化来提高训练速度,并使用域随机化来增强学习策略的鲁棒性。我们表明,与手工设计的奖励基线相比,STL塑形的奖励产生了更紧的速度跟踪和更稳定的训练。视频可在我们的项目网站上找到:this https URL。
英文摘要
Reinforcement learning (RL) for quadruped locomotion commonly depends on fixed, hand-crafted, and Markovian reward functions that may limit interpretability of learned policies and may lack explicit control over gait behaviors. We introduce a framework where distinct gaits are specified using parameterized constraints expressed in Signal Temporal Logic (STL). These include safety bounds, gait synchronization constraints, command tracking, and actuation bounds. From these specifications, we develop a reward shaping mechanism that provides learning agents a dense, continuous reward landscape that encodes desired behavior. We define parametric STL templates for three speed regimes (walking-trot, trot, bound), calibrate their parameters from reference rollouts, and compute rewards from using smooth approximations of STL robustness over the rollouts. The generated rewards can be used to provide shaped gradients compatible with Proximal Policy Optimization (PPO). We instantiate the approach on Google's Barkour quadruped robot in MuJoCo XLA (MJX). We use parallelization within the simulator to improve training speeds and use domain randomization to robustify learned policies. Compared with hand-crafted rewards, an expert-switching oracle, and Text2Reward, Human-STL maintains high command-tracking success across the evaluated speed range while exhibiting substantially higher consistency with the intended speed-dependent gait structures. Videos can be found on our project website: https://stl-locomotion.github.io/.
发表机构
- Department of Electrical Engineering and Computer Sciences University of California Berkeley(加州大学伯克利分校电气工程与计算机科学系)
- University of Southern California(南加州大学)
机构由 AI 辅助整理,请以论文原文为准。