发表机构
Autonomous Aerial Systems Lab; Technical University of Munich(自主空中系统实验室; 慕尼黑工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出一种结合残差动力学网络与学习型价值评价函数的MPC框架,利用GPU加速批处理iLQR求解器,在不完整动力学下提升闭环控制性能。
AI 中文摘要
模型预测控制(MPC)提供了一种结构化且具有约束意识的决策机制,但其对利于优化的解析动力学模型的依赖限制了其在涉及接触及其他难以建模的状态依赖任务中的应用。无模型强化学习避免了显式建模假设,但通常需要大量的交互数据。我们提出了一种基于学习的MPC框架,该框架结合了局部基于模型规划的数据效率和结构,以及用于补偿不完整动力学和有限视界短视的学习组件。该方法用残差动力学网络增强名义解析模型,该网络从数据中学习缺失的状态依赖效应,并将所得规划器与学习到的动作价值评价函数相结合,后者将长视界马尔可夫决策过程(MDP)结构注入局部迭代线性二次调节器(iLQR)优化中。为了使其在强化学习规模下切实可行,我们开发了一个GPU加速的批处理iLQR求解器,该求解器在最优控制回路内评估学习到的动力学和评价函数网络,并并行求解数千个轨迹优化问题。完整系统集成到机器人模拟器中,从而在不完整动力学下实现可扩展的基于模型的强化学习。在有偏且建模不完整的控制任务上的实验表明,该方法改善了闭环控制性能,同时保留了高效约束轨迹优化所需的基于模型的结构。
英文摘要
Model Predictive Control (MPC) provides a structured and constraint-aware mechanism for decision-making, but its reliance on optimization-friendly analytical dynamics models limits its use in tasks with contacts and other hard-to-model state dependencies. Model-free reinforcement learning avoids explicit modeling assumptions but typically requires large amounts of interaction data. We present a learning-based MPC framework that combines the data efficiency and structure of local model-based planning with learned components that compensate for incomplete dynamics and finite-horizon myopia. The method augments a nominal analytical model with a residual dynamics network that learns missing state-dependent effects from data and combines the resulting planner with a learned action-value critic that injects long-horizon MDP structure into the local iLQR optimization. To make this practical at reinforcement-learning scale, we develop a GPU-accelerated batched iLQR solver that evaluates learned dynamics and critic networks inside the optimal-control loop and solves thousands of trajectory-optimization problems in parallel. The complete system is integrated into a robotics simulator, enabling scalable model-based reinforcement learning under incomplete dynamics. Experiments on biased and incompletely modeled control tasks show that the approach improves closed-loop control performance while preserving the model-based structure needed for efficient constrained trajectory optimization.
CommentsKeywords: Model Predictive Control, Model-Based Reinforcement Learning, iLQR, Value Function Approximation. Presented at 10th Conference on Robot Learning (CoRL 2026), Austin TX, USA