AI 中文总结
本文提出PriEco-DRL框架,运用深度强化学习将电动公交生态驾驶与公交优先自适应信号控制相结合,采用优先级加权最大压力控制器及结构化奖励,经集中训练和分散执行,实验表明其能降低电动公交能耗,实现能量-时间权衡。
AI 中文摘要
城市交通电气化要求在电动公交(EB)的能源效率、时刻表可靠性和乘坐舒适性之间取得平衡,特别是在拥堵网络中与公交优先自适应信号交互时。本文提出了PriEco-DRL,这是一个联合优化框架,它使用深度强化学习(DRL)将EB生态驾驶与公交优先自适应信号控制集成在一起。信号层采用优先级加权最大压力(Priority-MP)控制器,根据占用感知压力分配绿灯时间,而车辆层根据不确定且动态变化的本地信号提示进行纵向控制。结构化奖励结合了引导和基于事件的强化,使EB到达与绿灯机会对齐,同时考虑能源、时间、舒适度和安全性。该框架使用具有参数共享的集中训练和分散执行(CTDE),允许单个DRL智能体使用本地观测从多辆公交车和路线学习。在实际走廊上的实验表明,PriEco-DRL与固定时间、感应和基于规则的信号-车辆协调基线相比,在保持网络效率和公交优先的同时降低了EB能耗。基于能量和轨迹的分析表明,改进源于自适应信号下更少的非计划启停事件和更平稳的速度调节。结果突出了可调的能量-时间权衡,通过奖励加权允许灵活的运营选择。
英文摘要
Urban transit electrification requires balancing energy efficiency, schedule reliability, and ride comfort for electric buses (EBs), particularly when interacting with transit-priority adaptive signals in congested networks. This paper proposes PriEco-DRL, a joint optimization framework that integrates EB eco-driving with transit-priority adaptive signal control using deep reinforcement learning (DRL). The signal layer employs a priority-weighted max-pressure (Priority-MP) controller to allocate green time based on occupancy-aware pressures, while the vehicle layer adapts longitudinal control based on uncertain and dynamically evolving local signal cues. A structured reward combines guidance and event-based reinforcement to align EB arrivals with green opportunities while considering energy, time, comfort, and safety. The framework uses centralized training and decentralized execution (CTDE) with parameter sharing, allowing a single DRL agent to learn from multiple buses and routes using local observations. Experiments on a real-world corridor show that PriEco-DRL reduces EB energy consumption while maintaining network efficiency and transit priority compared with fixed-time, actuated, and rule-based signal-vehicle coordination baselines. Energy- and trajectory-based analyses reveal that the improvements stem from fewer unscheduled stop-start events and smoother speed regulation under adaptive signals. The results highlight a tunable energy-time trade-off, allowing flexible operational choices through reward weighting.