发表机构
Delft University of Technology; Delft Center for Systems and Control(代尔夫特理工大学; 代尔夫特系统与控制中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对多类交通网络控制难题,提出DRL-MPC分层框架划分控制权,经多类高速公路网络实验,其性能优于混合状态反馈-MPC控制器,计算时间少于分层MPC控制器,模型失配时约束执行更有效。
AI 中文摘要
交通网络,尤其是多类交通网络(即包含混合车辆类型的网络),是复杂的控制难题。近年来,从与环境的交互中学习控制策略的深度强化学习(DRL),以及利用系统模型优化控制输入的模型预测控制(MPC),被越来越多地用于交通网络控制。然而,大规模网络中的非线性系统动力学和高维状态空间,在时间受限的训练下限制了DRL的学习能力,同时增加了MPC的计算时间,阻碍了其在计算资源有限场景下的实时部署。此外,MPC依赖于准确的网络模型,而多类交通网络这类复杂系统往往无法提供准确模型。本文提出了一种用于多类交通网络的新型DRL-MPC框架,该框架在DRL与MPC之间划分控制权,结合DRL的快速在线计算与模型无关性,以及MPC内置的优化与约束处理能力。在该分层框架中,MPC工作在高层,确定低频控制输入,其较慢的更新速率可适配其较高的计算时间;DRL工作在低层,利用其快速在线部署能力确定高频控制输入。该框架在多类高速公路网络上进行评估,对比了分层MPC控制器和混合状态反馈-MPC控制器,评估场景包含模型失配和有噪声的交通需求。结果表明,所提出的框架优于混合状态反馈-MPC控制器,与分层MPC控制器相比大幅减少了在线计算时间,且在模型失配情况下能实现更有效的约束执行。
英文摘要
Transportation networks, in particular multi-class transportation networks (i.e., networks with mixed vehicle types), are complex systems that are challenging to control. Recently, Deep Reinforcement Learning (DRL), which learns control policies from interactions with the environment, and Model Predictive Control (MPC), which uses a system model to optimize control inputs, have been increasingly utilized for transportation network control. However, nonlinear system dynamics and high-dimensional state spaces in large-scale networks limit DRL's learning capacity under time-constrained training and increase MPC's computation time, hindering real-time implementation with limited computational resources. Moreover, MPC depends on an accurate network model, which is often unavailable for complex systems such as multi-class transportation networks. This paper proposes a novel DRL-MPC framework for multi-class transportation networks that divides control authority between DRL and MPC, combining DRL's fast online computation and model independence with MPC's built-in optimization and constraint-handling capabilities. In the hierarchical framework, MPC operates at the higher level and determines low-frequency control inputs whose slower update rate accommodates its high computation time, while DRL operates at the lower level and determines high-frequency control inputs using its fast online deployment. The framework is evaluated on a multi-class freeway network against a hierarchical MPC controller and a hybrid state-feedback-MPC controller, including scenarios with model mismatch and noisy traffic demands. Results show that the proposed framework outperforms the hybrid state-feedback-MPC controller, substantially reduces online computation time compared with the hierarchical MPC controller, and provides more effective constraint enforcement under model mismatch.