arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SKooP:用于强化学习的更快、更通用的有腿机器人运动的对称库普曼预测

SKooP: Symmetric Koopman Predictions for Faster and More Generalizable Legged Robot Locomotion with Reinforcement Learning

Evelyn D'Elia, Weishu Zhan, Giulio Turrisi, Giulio Romualdi, Giuseppe L'Erario, Raffaello Camoriano, Wei Pan, Daniele Pucci

arXiv 2607.11624首次发表:更新:

发表机构

Italian Institute of Technology (IIT); University of Manchester; Dynamic Legged Systems Laboratory, IIT; Generative Bionics S.R.L; DAUIN, Politecnico di Torino; Rehab Technologies Lab, IIT; Newcastle University(意大利理工学院; 曼彻斯特大学; 意大利理工学院动态腿式系统实验室; 生成仿生学有限公司; 都灵理工大学DAUIN; 意大利理工学院康复技术实验室; 纽卡斯尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对强化学习样本效率低问题展开,提出SKooP方法,结合形态对称与库普曼模型优势,在学习策略时学习系统动力学模型,将其预测用于评论家,还纳入群对称,验证了该方法能减少收敛时间并增加学习奖励,且策略可转移。

AI 中文摘要

强化学习算法通常样本效率较低。在机器人领域,最近一系列工作通过在学习过程中编码物理先验来解决此问题,但大多在定义明确的低维基准系统上验证,而非复杂非线性动力学的高维机器人。本文介绍了SKooP(对称库普曼预测),它结合形态对称优势与通过自动编码器学习的库普曼模型优势来增强策略学习。SKooP在学习策略时同时学习系统动力学的库普曼模型,其预测用作评论家的特权观测,还将群对称纳入网络以产生高度等变策略。通过深入分析学习到的库普曼模型和对称策略验证了SKooP,展示了它们如何影响智能体性能,还表明学习到的策略可转移到不同模拟环境。结果表明SKooP持续减少收敛时间并增加四足机器人上多个具有挑战性的双足运动任务的学习奖励。

英文摘要

Reinforcement learning (RL) algorithms classically suffer from poor sample efficiency. In robotics, a recent line of work has emerged addressing this problem by encoding physics priors in the learning process. However, most of these approaches are validated on well-defined, low-dimensional benchmark systems rather than high-dimensional robots with complex nonlinear dynamics. In this paper, we introduce \textit{SKooP (Symmetric Koopman Predictions)}, an approach combining the advantages of morphological symmetries with those of a Koopman model learned via autoencoder to enhance policy learning. SKooP learns a Koopman model of the system dynamics alongside the policy. The resulting Koopman predictions are used as privileged observations for the critic, allowing the agent to learn based on smoother, more informative features. We also incorporate group symmetries into the actor, critic, encoder and decoder networks to produce a highly equivariant policy. The SKooP approach is validated via in-depth analysis of the learned Koopman models and symmetric policies to showcase how each of these influences the agent's performance. We also show that the learned policies are transferable to different simulation environments. Our results show that SKooP consistently reduces convergence time and increases the learned reward for multiple challenging bipedal locomotion tasks on a quadruped robot. Project page: https://evelyd.github.io/SymmetricKoopmanPredictions

CommentsThis paper has been accepted for publication at the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Pittsburgh, USA, 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑