发表机构
University of Groningen(格罗宁根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出首个使用可微仿真训练可部署盲四足机器人运动策略的开源框架Open-DiffLoco,该框架实现SHAC算法并提出JAVE扩展,训练的策略在Unitree Go2机器人上表现优异且资源消耗低。
AI 中文摘要
通过传统强化学习开发可部署的运动策略通常需要复杂的奖励工程和高昂的训练时间。虽然可微仿真提供了一种高效的替代方案,但能够将这些策略端到端迁移到物理硬件的开源工具仍然有限。本文提出Open-DiffLoco,一种用于通过可微仿真训练可部署盲四足机器人运动策略的开源框架。该框架在MuJoCo XLA(MJX)中实现了Short-Horizon Actor-Critic(SHAC)算法,并训练了可迁移到真实硬件的本体感知策略。部署的策略移除了基线性速度等特权演员观测,且不依赖参考轨迹,还使用了大幅简化的奖励函数,使机器人无需传统强化学习流程中典型的复杂辅助奖励即可发现行走模式。当部署在Unitree Go2四足机器人硬件上时,训练后的策略跟踪全向速度指令的均方根误差低于0.2 m/s,速度可达1 m/s以上,且对不平地形和侧向推等外部物理干扰具有鲁棒性。在报告的所有配置中,训练在单个NVIDIA GeForce RTX 5080 GPU上使用不足6 GB的VRAM,耗时约20-60分钟。作为SHAC的算法扩展,本文提出Jacobian-Augmented Value Estimation(JAVE),通过监督评论家雅可比矩阵来改进早期一阶策略梯度训练。据作者所知,Open-DiffLoco是首个使用可微仿真训练可部署运动策略的开源框架,部署视频和源代码可在指定URL获取。
英文摘要
Developing deployable locomotion policies through conventional reinforcement learning often requires complex reward engineering and expensive training times. While differentiable simulation offers a highly efficient alternative, open-source tools capable of end-to-end transfer of these policies to physical hardware remain limited. This paper introduces Open-DiffLoco, an open-source framework for training deployable blind quadruped locomotion policies with differentiable simulation. The framework implements the Short-Horizon Actor-Critic (SHAC) algorithm in MuJoCo XLA (MJX) and trains a proprioceptive policy that transfers to real-world hardware. The deployed policy removes privileged actor observations, including base linear velocity, and does not rely on reference trajectories. It also uses a substantially simplified reward function, enabling the robot to discover walking patterns without the complex auxiliary rewards typically used in conventional reinforcement learning pipelines. When deployed on physical hardware (a Unitree Go2 quadruped), the trained policy tracks omnidirectional velocity commands with root-mean-square error below 0.2 m/s, reaches speeds above 1 m/s, and remains robust to uneven terrain and external physical disturbances, such as lateral pushes. Across the reported configurations, training uses under 6 GB of VRAM on a single NVIDIA GeForce RTX 5080 GPU and completes in approximately 20-60 minutes. As an algorithmic extension to SHAC, we propose Jacobian-Augmented Value Estimation (JAVE), which supervises the critic Jacobians to improve early first-order policy-gradient training. To our knowledge, Open-DiffLoco is the first open-source framework for training deployable locomotion policies using differentiable simulation. Deployment videos and source code are available at: https://diffloco.martin-opat.com/
Comments8 pages, 5 figures, Project page, videos, and code available at: https://diffloco.martin-opat.com/