arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26860cs.LGcs.AImath.OC

基于强化学习的混合交通环境下网联自动驾驶车辆(CAV)队列汇入机动控制

Reinforcement Learning-Based Control of CAV Platoon Joining Maneuvers in Mixed Traffic

  • Université Gustave Eiffel(古斯塔夫·埃菲尔大学)
  • Université Paris Dauphine-PSL(巴黎多芬大学-PSL)
  • Cosys-Grettia, Univ Gustave Eiffel(古斯塔夫·埃菲尔大学Cosys-Grettia机构)

机构由 AI 辅助整理,请以论文原文为准。

Biao Yin, Abderrahmane Kasmi, Nadir Farhi

AI总结:

该研究针对混合交通环境下CAV队列汇入的安全高效控制问题,提出结合SUMO的建模框架,评估PPO等DRL算法,发现PPO在平衡安全与效率上表现更优。

AI中文摘要:

网联自动驾驶车辆(CAV)队列行驶为提升道路安全性和交通通行能力提供了极具前景的方案,但现实交通中的队列控制因不确定性和异质驾驶行为而颇具挑战。强化学习(RL)在解决此类控制问题上潜力巨大,但其实际部署面临安全性和学习效率相关的挑战。本文提出一种通用建模与仿真框架,用于研究CAV队列汇入机动并比较基于深度强化学习(DRL)的控制算法。该问题在混合交通环境中尤为棘手,CAV与表现出异质纵向和横向行为的人类驾驶车辆共存。研究目标是通过两种方式实现安全高效的汇入机动:一是在学习过程中纳入危险行为的惩罚项,二是使用外部安全控制器约束学习到的策略。本文采用基于智能体的建模框架与城市移动仿真(SUMO)模拟器相结合的方式,对深度Q网络(DQN)、双深度Q网络(DDQN)和近端策略优化(PPO)算法进行评估。结果显示,PPO的性能优于DQN和DDQN,其汇入成功率约为98%,碰撞率低于1%,这主要归功于奖励函数中纳入的风险相关惩罚项;不过,这种性能提升需要更多决策步骤来完成机动,揭示了安全性、汇入效果与决策效率之间的权衡关系。外部安全控制器可有效防止碰撞,但其干预可能会降低汇入效率。研究结果强调,在为混合交通环境中的CAV队列汇入设计基于RL的控制器时,需同时考虑安全性与效率。

英文摘要:

Connected and automated vehicle (CAV) platooning offers a promising approach to improving road safety and traffic capacity. However, platoon control in real-world traffic is challenging due to uncertainty and heterogeneous driving behaviors. Reinforcement learning (RL) has strong potential for addressing such control problems, but its practical deployment raises challenges related to safety and learning efficiency. This paper proposes a generic modeling and simulation framework for investigating CAV platoon joining maneuvers and comparing deep reinforcement learning (DRL)-based control algorithms. The problem is particularly challenging in mixed-traffic environments, where CAVs coexist with human-driven vehicles exhibiting heterogeneous longitudinal and lateral behaviors. The objective is to achieve safe and efficient joining maneuvers by either incorporating penalties for risky behaviors into the learning process or using an external safety controller to constrain the learned policy. An agent-based modeling framework coupled with the Simulation of Urban MObility (SUMO) simulator is used to evaluate Deep Q-Network (DQN), Double Deep Q-Network (DDQN), and Proximal Policy Optimization (PPO). Results show that PPO outperforms DQN and DDQN, achieving a joining success rate of approximately 98 % and a collision rate below 1 %, largely due to risk-related penalties incorporated into the reward function. However, this improved performance requires more decision steps to complete the maneuver, revealing a trade-off between safety, joining effectiveness, and decision efficiency. An external safety controller effectively prevents collisions, although its interventions may reduce joining efficiency. The results highlight the importance of jointly considering safety and efficiency when designing RL-based controllers for CAV platoon joining in mixed traffic.

↑