发表机构
Delft University of Technology(代尔夫特理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出复合梯度学习(CGL)方法,将MPC整合进DRL训练过程,通过联合动作表示和交互考虑,在强交互场景下提升控制策略性能,但平均增益有限。
AI 中文摘要
集成深度强化学习(DRL)和模型预测控制(MPC)的方法越来越多地用于控制自主系统,通过结合它们的互补能力。DRL通过与环境的交互来学习控制策略。MPC使用系统模型来优化控制输入,同时考虑约束条件。在具有共享控制权限的DRL-MPC框架中,DRL智能体和MPC控制器各自确定部分控制输入。然而,常见的学习公式将MPC视为环境的一部分,因此没有明确考虑MPC对控制的贡献或其与DRL智能体的交互。本文提出了一种新颖的复合梯度学习(CGL)方法,通过将DRL和MPC控制输入表示为联合动作,并在训练期间更新DRL智能体时考虑它们的交互,从而将MPC控制器整合到学习过程中。CGL在两个具有不同DRL和MPC控制输入交互强度的多类高速公路交通网络上进行了评估,并与将MPC视为环境一部分或仅部分将MPC纳入学习的替代方法进行了比较。结果表明,在弱交互下,CGL提供的益处有限,但在强交互下,CGL在部分训练运行中学习到了比替代方法性能更高的控制策略,尽管平均控制性能提升仍然有限。
英文摘要
Integrated deep reinforcement learning (DRL) and model predictive control (MPC) methods are increasingly used to control autonomous systems by combining their complementary capabilities. DRL learns control policies through interaction with the environment. MPC uses a system model to optimize control inputs while accounting for constraints. In DRL-MPC frameworks with shared control authority, both the DRL agent and the MPC controller each determine part of the control inputs. However, common learning formulations treat MPC as part of the environment and therefore do not explicitly account for MPC's contribution to control or its interaction with the DRL agent. This paper proposes a novel composite-gradient learning (CGL) method that integrates the MPC controller into the learning process by representing the DRL and MPC control inputs as a joint action and accounting for their interaction when updating the DRL agent during training. CGL is evaluated on two multi-class freeway traffic networks with different strengths of interaction between the DRL and MPC control inputs and it is compared with alternative methods that treat MPC as part of the environment or that only partially incorporate MPC into learning. The results show that CGL offers limited benefit under weak interaction, but learns higher-performing control policies than the alternative methods in a subset of training runs under strong interaction, although the average control-performance gains remain modest.