从LLM生成的规范到学习四足运动
From LLM-Generated Specifications to Learned Quadruped Locomotion
浏览论文内容
中文总结 AI 辅助
本研究利用LLM生成PSTL规范,转化为奖励函数训练四足运动策略,在MJX中验证,步态感知Qwen 3.6规范在高速下优于Text2Reward。
中文摘要 AI 辅助
四足机器人运动策略通常使用强化学习进行训练,而强化学习又严重依赖手工设计的奖励函数。设计奖励函数需要大量的手动工程,并且通常不清楚哪些局部奖励会引发所需的全局行为。来自形式规范(如信号时序逻辑(STL))的塑形奖励可以使奖励更具可解释性,但编写STL规范本身仍需要领域专业知识。我们研究大型语言模型(LLM)是否可以通过生成参数化信号时序逻辑(PSTL)规范来填补这一空白,这些规范随后用于策略学习。给定自然语言运动目标和受约束的规范语法,GPT-5.5和Qwen 3.6独立提出用于命令跟踪、安全性和步态结构的STL模板。我们使用专家轨迹实例化生成的PSTL模板的参数,并仅保留与演示的专家行为一致的规范。然后将所得规范转换为平滑的有限历史奖励函数,并使用近端策略优化(PPO)在MuJoCo XLA(MJX)中训练四足运动策略。我们评估了“步态感知”和“步态无关”两种设置。前者指定行走-快步、快步和跳跃模式,而后者允许接触模式从任务目标中涌现。我们与手工设计的奖励、Text2Reward风格的LLM生成的奖励代码以及专家切换预言机进行比较。步态感知的Qwen 3.6规范在所有测试速度(0.3--2.1 m/s)下实现了100%的存活率和命令成功率,并在高速下匹配目标步态,而Text2Reward在≥1.9 m/s时两个指标均为0%。视频:此https URL
英文摘要
Quadruped robot locomotion policies are often trained using reinforcement learning, which in turn relies heavily on hand-crafted reward functions. Designing reward functions requires substantial manual engineering, and it is often unclear which local rewards will induce the desired global behavior. Shaped rewards from formal specifications in languages like Signal Temporal Logic (STL) can make rewards more interpretable, but writing STL specifications itself still requires domain expertise. We study whether large language models (LLMs) can fill this gap by generating Parametric Signal Temporal Logic (PSTL) specifications that are subsequently used for policy learning. Given a natural language locomotion objective and a constrained specification grammar, GPT-5.5 and Qwen 3.6 independently propose STL templates for command tracking, safety, and gait structure. We instantiate the parameters of the generated PSTL templates using expert trajectories and retain only specifications that are consistent with demonstrated expert behavior. The resulting specifications are then transformed into smooth, finite-history reward functions and used to train a quadruped locomotion policy with Proximal Policy Optimization (PPO) in MuJoCo XLA (MJX). We evaluate both \emph{gait-aware} and \emph{gait-agnostic} settings. The former specifies walking-trot, trot, and bound regimes, while the latter allows contact patterns to emerge from the task objective. We compare against hand-engineered rewards, Text2Reward-style LLM-generated reward code, and an expert-switching oracle. Gait-aware Qwen 3.6 specifications achieved 100\% survival and command success across all tested speeds (0.3--2.1 m/s) and matched the target gait at high speeds, whereas Text2Reward achieved 0\% for both metrics at $\geq 1.9$ m/s. Videos: https://stl-locomotion.github.io/
发表机构
- University of Southern California(南加州大学)
- University of Florida(佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。