arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35152math.OCcs.LG

连续交织路段考虑合流/分流风险的协调车道级可变限速与匝道控制:一种混合模型预测控制与多智能体强化学习方法

Coordinated Lane-Level Variable Speed Limits and Ramp Metering for Successive Weaving Segments Considering Merging/Diverging Risks: A Hybrid Model Predictive Control and Multi-Agent Reinforcement Learning Approach

Guodong Ma, Baofeng Sun, Wenyu Yang, Zhihong Yao

首次发表
浏览论文内容

中文总结 AI 辅助

针对城市快速路连续交织路段,提出融合车道级宏观模型与MPC-MARL分层控制的混合框架,协调可变限速与匝道控制,降低合流/分流风险,实测验证其精度与泛化性优于现有方法。

中文摘要 AI 辅助

城市快速路上的连续交织路段(SWSs)是易发生周期性拥堵和碰撞的瓶颈区域,需要精细化的主动交通管理(ATM)。现有方法难以平衡数据驱动优化的自适应性能与基于模型控制的鲁棒性和可迁移性。我们提出了一种混合框架,用于协调跨连续交织路段的车道级可变限速(VSLs)和匝道控制。首先,我们重构了L-METANET,一种能够捕捉自由和强制车道变换的车道级宏观交通流模型。其次,我们将XGBoost-SHAP与随机参数二元Logit(RPBL)模型相结合,推导出合流和分流碰撞风险的分析方程,并构建了系统成本和奖励函数。第三,我们开发了MPC-STMAPPO,一种集成模型预测控制(MPC)和多智能体强化学习(MARL)的分层控制器。其上层MPC层利用L-METANET进行长时域滚动优化并生成基线指令;其下层时空MAPPO(ST-MAPPO)层,通过Mamba单元和图注意力增强,产生用于短时域调整的残差动作。在中国长春18公里东部快速路上的真实世界实验表明,L-METANET准确再现了由车道变换引起的流量重新分布和通行能力下降,状态演化与地面真实数据一致。XGBoost-SHAP-RPBL在大多数任务中AUC超过0.80,优于传统Logit模型。MPC-STMAPPO在多个指标上比基于MPC和MARL的基线收敛更快且性能更好。在随机波动需求下,它在泛化方面也显著优于纯MARL,展现出强大的工业部署潜力。

英文摘要

Successive weaving segments (SWSs) on urban expressways are bottlenecks prone to recurrent congestion and collisions, requiring fine-grained active traffic management (ATM). Existing approaches struggle to balance the adaptive performance of data-driven optimization with the resilience and transferability of model-based control. We propose a hybrid framework to coordinate lane-level variable speed limits (VSLs) and ramp metering across SWSs. First, we reconstruct L-METANET, a lane-level macroscopic traffic flow model that captures free and forced lane changes. Second, we combine XGBoost-SHAP with a random-parameters binary logit (RPBL) model to derive analytical equations for merging and diverging collision risks and formulate system cost and reward functions. Third, we develop MPC-STMAPPO, a hierarchical controller integrating model predictive control (MPC) and multi-agent reinforcement learning (MARL). Its upper MPC layer uses L-METANET for long-horizon rolling optimization and generates baseline commands; its lower spatiotemporal MAPPO (ST-MAPPO) layer, enhanced with Mamba cells and graph attention, produces residual actions for short-horizon adjustment. Real-world experiments on the 18-km Eastern Expressway in Changchun, China, show that L-METANET accurately reproduces lane-changing-induced flow redistribution and capacity drops, with state evolution aligned with ground truth. XGBoost-SHAP-RPBL achieves AUCs above 0.80 in most tasks, outperforming conventional logit models. MPC-STMAPPO converges faster and performs better across multiple metrics than MPC- and MARL-based baselines. Under randomly fluctuating demand, it also significantly outperforms pure MARL in generalization, demonstrating strong potential for industrial deployment.

发表机构

  • Jilin University(吉林大学)
  • Southwest Jiaotong University(西南交通大学)

机构由 AI 辅助整理,请以论文原文为准。

↑