arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23204cs.RO

引导式黎曼优化(GuRO):连接模型预测控制与决策Transformer

Guided Riemannian Optimization (GuRO): Bridging Model Predictive Control and Decision Transformers

  • The University of Manchester(曼彻斯特大学)

机构由 AI 辅助整理,请以论文原文为准。

Hossein Abdi, Satya Prakash Dash, Mingfei Sun

AI总结:

本研究提出GuRO框架,将MPC与RL整合到序列决策框架,利用黎曼优化解决非凸损失问题,在四足机器人控制任务上优于TRPO等基准方法。

AI中文摘要:

高维非线性系统中的决策是机器人领域的核心挑战。模型预测控制(MPC)这类基于模型的方法具备样本效率和可解释性,但在动力学模型不准确或需要长 horizon 预测时性能会下降;而无模型强化学习(RL)直接从交互中学习策略,却存在样本复杂度高、优化不稳定的问题。近期序列建模的进展催生了基于Transformer的决策框架,可统一MPC与RL,但这类框架的训练通常因高度非凸的损失景观面临显著优化挑战。本研究提出一种新型框架,将MPC与RL整合到序列决策框架中,并利用曲率感知优化高效应对非凸损失景观。MPC提供局部最优轨迹的预测以引导决策Transformer,无需大量离线预训练;为解决传统优化器收敛慢、不稳定的问题,我们采用高效黎曼(曲率感知)方法在黎曼参数空间训练策略,实现更快、更鲁棒的优化。我们在高维四足机器人控制任务上评估该框架,结果显示其在包括TRPO、SAC和在线决策Transformer在内的强基准方法上实现了一致的性能提升,获得更高回报且收敛更快。

英文摘要:

Decision-making in high-dimensional, nonlinear systems remains a central challenge in robotics. While model-based methods like Model Predictive Control (MPC) offer sample efficiency and interpretability, their performance degrades when the dynamics model is inaccurate or long-horizon predictions are required. Conversely, model-free reinforcement learning (RL) learns policies directly from interaction but suffers from high sample complexity and unstable optimization. Recent advances in sequence modeling have inspired transformer-based decision-making frameworks that can unify MPC and RL, but their training typically faces significant optimization challenges due to highly non-convex loss landscapes. In this work, we propose a novel framework that integrates MPC with RL in a sequence decision-making framework and leverages a curvature-aware optimization to efficiently tackle non-convex loss landscapes. MPC provides predictions of locally optimal trajectories that guide the decision transformer, removing the need for extensive offline pretraining. To address the slow and unstable convergence of traditional optimizers, we train the policy in a Riemannian parameter space using an efficient Riemannian (curvature-aware) method, leading to faster and more robust optimization. We evaluate our framework on high-dimensional quadruped control tasks and demonstrate consistent improvements over strong baselines, including TRPO, SAC, and Online Decision Transformer, achieving higher returns and faster convergence.

↑