AI 中文总结
研究针对多智能体强化学习在交通信号控制中应对不利情况能力不足的问题,提出集成自适应上下文博弈最坏情况估计器的分布鲁棒框架,经实验验证其能提升鲁棒性和效率,且策略有零样本泛化能力。
AI 中文摘要
多智能体强化学习已成为交通信号控制的一种有前景的方法。然而,标准的多智能体强化学习策略通常在名义条件下针对预期回报进行优化,在不利情况下极易受到时空需求变化和灾难性拥堵的影响。为解决这一关键限制,本文提出了一种与算法无关的分布鲁棒多智能体强化学习框架,集成了自适应上下文博弈最坏情况估计器。在较慢的时间尺度上运行,该估计器在训练期间通过动态生成对抗性需求混合与交通控制器共同进化。这引导学习过程强化策略以抵御瓶颈场景,而无需修改底层的多智能体强化学习架构。该框架在合成的5x5网格和异构的摩纳哥城市网络上,针对基于价值、演员-评论家以及策略梯度方法进行了评估。实证结果表明,该分布鲁棒框架可防止队列无界增长,并显著提高最坏情况的鲁棒性和平均情况的效率。对于摩纳哥环境中的近端策略优化架构,平均而言,鲁棒再训练使最坏情况队列长度减少了74.39%,并使全网络平均情况队列长度提高了75.45%。此外,再训练后的策略对未见的交通分布表现出强大的零样本泛化能力,突出了该框架的可扩展性以及在弹性现实世界城市部署中的潜力。
英文摘要
Multi-agent reinforcement learning (MARL) has emerged as a promising approach for traffic signal control. However, standard MARL policies typically optimize for expected returns under nominal conditions, leaving them highly vulnerable to spatial-temporal demand shifts and catastrophic congestion under adverse scenarios. To address this critical limitation, this paper proposes an algorithm-agnostic Distributionally Robust (DR) MARL framework integrating an adaptive Contextual-Bandit Worst-Case Estimator (CB-WCE). Operating on a slower timescale, the CB-WCE co-evolves with the traffic controllers by dynamically generating adversarial demand mixtures during training. This steers the learning process to fortify policies against bottleneck scenarios without requiring modifications to the underlying MARL architectures. The framework is evaluated across value-based, actor-critic, and policy-gradient methods on both a synthetic 5x5 grid and a heterogeneous Monaco City network. Empirical results demonstrate that the DR framework prevents unbounded queue growth and profoundly enhances both worst-case robustness and average-case efficiency. Notably, for the Proximal Policy Optimization (PPO) architecture in the Monaco environment, on average, robust retraining reduced the worst-case queue length by 74.39% and improved the average-case network-wide queue length by 75.45%. Furthermore, the retrained policies exhibit strong zero-shot generalization to unseen traffic distributions, highlighting the framework's scalability and potential for resilient real-world urban deployment.