SelectLight:学习选择分布式模型预测控制生成的城市交通网络信号方案
SelectLight: Learning to Select Signal Plans Generated by Distributed Model Predictive Control for Urban Traffic Networks
浏览论文内容
中文总结 AI 辅助
SelectLight通过MARL策略选择DMPC生成的交通信号方案,在SUMO网络实验中延迟性能最优,需求加倍时排队延迟降低5.57%,求解时间远低于控制间隔。
中文摘要 AI 辅助
城市交通网络的协调交通信号控制必须适应不断变化的需求,同时满足操作约束。多目标分布式模型预测控制(DMPC)可在线构建可行的信号方案,但用于在权衡解中进行选择的规定规则无法从已实现的闭环结果中学习。我们提出SelectLight,该方法通过允许多智能体强化学习(MARL)策略直接从DMPC在线生成的方案中进行选择,实现了优化后选择。在每次控制更新时,状态剪枝多目标动态规划(SP-MODP)使用Newellian点-空间队列模型评估方案,并返回一组有界的、相互非支配的候选信号方案,涉及总排队延迟、峰值队列积累和总停车次数三个目标。采用独立近端策略优化(IPPO)训练的拓扑感知注意力策略,从每个可变大小的集合中选择一个未修改的方案。这将学习限制在候选选择上,保留规定的信号定时约束,并使所选方案及其预测的目标权衡可供检查。在两个28交叉口的SUMO网络上进行的实验表明,SelectLight实现了最佳的延迟相关性能,且其优势随需求增加而扩大。在两倍基线需求下,与最强基线相比,它分别减少了5.57%的排队延迟和6.44%的等待时间。SelectLight在所有测试的需求变化下也产生最低的转移损失。在120秒的预测时域下,每个交叉口的第99百分位SP-MODP求解时间为5.408毫秒,远低于5秒的控制间隔。
英文摘要
Coordinated traffic signal control across urban networks must adapt to changing demand while satisfying operational constraints. Multi-objective distributed model predictive control (DMPC) can construct feasible signal plans online, but prescribed rules for selecting among trade-off solutions cannot learn from realized closed-loop outcomes. We propose SelectLight, which implements post-optimization selection by allowing a multi-agent reinforcement learning (MARL) policy to choose directly from plans generated online by DMPC. At each control update, state-pruned multi-objective dynamic programming (SP-MODP) evaluates plans with a Newellian point--spatial queue model and returns a bounded set of mutually nondominated candidate signal plans for total queueing delay, peak queue accumulation, and total number of stops. A topology-aware attention policy trained with independent proximal policy optimization (IPPO) selects one unmodified plan from each variable-size set. This confines learning to candidate selection, preserves the prescribed signal timing constraints, and leaves the selected plan and its predicted objective trade-offs available for inspection. Experiments on two 28-intersection SUMO networks show that SelectLight achieves the best delay-related performance and that its advantage widens with demand. At twice the baseline demand, it reduces queueing delay and waiting time by 5.57% and 6.44%, respectively, relative to the strongest baseline. SelectLight also incurs the lowest transfer loss under every tested demand shift. With a 120 s prediction horizon, the per-intersection 99th-percentile SP-MODP solution time is 5.408 ms, well below the 5 s control interval.