发表机构
The University of Manchester; BAE Systems; University of Warwick(曼彻斯特大学; BAE系统公司; 华威大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出PROMO,一种偏好条件多目标强化学习方法,将偏好作为运行时输入,使单一策略实现多目标权衡,在仿真和真实机器人上展现广泛帕累托覆盖和显著性能提升。
AI 中文摘要
四足运动需要在指令跟踪、稳定性和能量效率等相互冲突的目标之间进行平衡,然而传统强化学习(RL)在训练时将优先级硬编码为固定的标量奖励。我们提出PROMO(偏好条件多目标强化学习),这是一种语义多目标方法,将这种权衡作为单一运动策略的显式运行时输入。PROMO在保持具体形态的运动先验不变的同时,根据部署时面对的偏好来条件化策略,从而将操作者意图与生成可行步态所需的奖励塑形项分离开来。与固定目标控制器、多目标基线和独立训练的专业模型相比,PROMO从单一可部署策略中实现了目标专业化和鲁棒性。在仿真中采样的100个偏好下,67个行为在精确帕累托支配下是非支配的,平均偏好-目标相关性为0.843,展示了广泛的帕累托覆盖和可预测的偏好响应。同一策略零样本迁移到Unitree Go2,与平衡偏好相比,仅偏好变化就使比能耗降低高达30.4%,位置误差降低38.7%,峰值身体姿态偏差降低59.0%。这些结果确立了偏好条件多目标强化学习作为自适应腿式运动的实用运行时接口,将其作用扩展到离线帕累托集构建之外。开源代码和视频可在该https URL获取。
英文摘要
Quadrupedal locomotion requires balancing conflicting objectives such as command tracking, stability, and energy efficiency, yet conventional reinforcement learning (RL) hardcodes these priorities into a fixed scalar reward at training time. We present PROMO (Preference-Conditioned Multi-Objective Reinforcement Learning), a semantic multi-objective approach that makes this trade-off an explicit runtime input to a single locomotion policy. PROMO conditions the policy on deployment facing preferences while keeping embodiment-specific locomotion priors fixed, thereby separating operator intent from reward shaping terms required for viable gait generation. Compared with fixed-objective controllers, multi-objective baselines, and independently trained specialists, PROMO achieves objective specialization and robustness from a single deployable policy. Across 100 sampled preferences in simulation, 67 behaviors are non-dominated under exact Pareto dominance, with a mean preference-objective correlation of 0.843, demonstrating broad Pareto coverage and predictable preference response. The same policy transfers zero-shot to a Unitree Go2, where preference changes alone reduce specific energy by up to 30.4%, position error by 38.7%, and peak body-attitude deviation by 59.0% relative to the balanced preference. These results establish preference-conditioned multi-objective RL as a practical runtime interface for adaptive legged locomotion, extending its role beyond offline Pareto-set construction. Open-source code and videos are available at https://amrmousa.com/promo/.
CommentsSubmitted to IEEE Transactions on Robotics. Project website, code, and videos: https://amrmousa.com/promo/