arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过针对权重变化的模型预测控制(MPC)的MPC求解器梯度引导加速强化学习

Accelerating Reinforcement Learning via MPC Solver-Gradient Guidance for Weights-varying MPC

Baha Zarrouki, Arslan Thobani, Jasper Hoffmann, Mattia Piccinini, Rudolf Reiter, Felix Jahncke, Sébastien Gros, Davide Scaramuzza, Johannes Betz

arXiv 2609.01061首次发表:更新:

AI 中文总结

该研究提出SG-RL算法,将MPC求解器梯度作为辅助引导,在PPO中实现MPC代价权重自适应,在自主赛车平台上样本效率提升70.6%,性能优于GB-PL且可零样本泛化。

AI 中文摘要

在模型预测控制(MPC)中,代价函数权重决定闭环行为,但工况变化常使固定参数化方案不再最优,促使需依赖上下文的在线自适应。学习此类策略颇具挑战,因为行为隐含依赖于数值MPC求解结果,会产生对策略参数的非线性、可能非光滑的长程依赖,进而产生偏差-方差权衡:强化学习(RL)从环境样本中优化实际闭环回报,但样本效率低;而基于梯度的策略学习(GB-PL)利用可微分MPC的低方差求解器梯度,在预测轨迹上优化代理损失,但在模型失配时可能存在偏差。我们提出求解器梯度引导强化学习(SG-RL),这是一种针对基于RL的在线MPC代价权重自适应的求解器灵敏度增强方法。SG-RL将采样的闭环回报作为目标,并利用有界的求解器衍生梯度作为辅助引导,以提升稳定性和样本效率。我们在近端策略优化(PPO)中实例化SG-RL,开发了四种模块化算法,分别将求解器梯度引导注入演员更新缩放、策略损失、优势估计和价值函数学习。在存在故意模型失配的两个全尺寸自主赛车平台上,SG-RL达到PPO的最佳闭环回报,样本量减少多达70.6%,闭环回报比GB-PL基线至少高出54%,且能零样本泛化到未见环境。

英文摘要

In Model Predictive Control (MPC), cost-function weights shape closed-loop behavior, yet changing conditions often make fixed parametrizations suboptimal and motivate context-dependent online adaptation. Learning such policies is difficult because behavior depends implicitly on numerical MPC solutions, producing nonlinear, potentially nonsmooth, long-horizon dependencies on policy parameters. This creates a bias-variance tradeoff: Reinforcement Learning (RL) optimizes realized closed-loop return from environment samples but is sample-inefficient, whereas Gradient-Based Policy Learning (GB-PL) uses low-variance solver gradients from differentiable MPC to optimize surrogate losses on predicted trajectories but can be biased under model mismatch. We propose Solver-Gradient Guided Reinforcement Learning (SG-RL), a solver-sensitivity augmentation for RL-based online MPC cost-weight adaptation. SG-RL keeps sampled closed-loop return as the objective and uses bounded solver-derived gradients as auxiliary guidance to improve stability and sample efficiency. We instantiate SG-RL in Proximal Policy Optimization (PPO) with four modular algorithms that inject solver-gradient guidance into actor-update scaling, policy loss, advantage estimation, and value-function learning. On two full-scale autonomous racing platforms with intentional model mismatch, SG-RL reaches PPO's best closed-loop return with up to 70.6% fewer samples, outperforms GB-PL baselines by at least 54% in closed-loop return, and generalizes zero-shot to unseen environments.

CommentsSubmitted to IEEE Transactions on Robotics (T-RO). 18 pages, 13 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑