采用考虑路段流量传播引导的强化学习方法的动态OD矩阵在线估计
Online Estimation of Dynamic Origin-Destination Matrices Using Reinforcement Learning with Link-Flow Propagation Guidance
浏览论文内容
中文总结 AI 辅助
本研究提出LFPG-RL方法,将路段流量传播引导集成到近端策略优化中,在墨尔本干线网络测试中实现了更高效准确的在线动态OD矩阵估计。
中文摘要 AI 辅助
动态起讫点(OD)矩阵在线估计(DODE)用于校准随时间变化的OD需求,以复现观测到的路段流量轨迹。在线场景下,需根据当前观测值和传播后的网络状态估计OD需求,而后续观测值与随机动态网络加载(DNL)结果仍存在不确定性。近年来,强化学习(RL)成为颇具潜力的替代方案,它通过替代迭代算法降低计算负担,且适用于随机环境。但由于策略是离线训练后在线部署的,需处理不同的目标路段流量轨迹;每个目标轨迹定义了奖励中使用的路段流量误差,因此同一OD需求向量可能需要不同调整,导致传统标量反馈存在歧义。为解决这一问题,本研究提出LFPG-RL,将路段流量传播引导(LFPG)集成到近端策略优化(PPO)中。LFPG结合路段流量误差敏感性与各OD-时间需求分量对模拟路段流量的贡献,将聚合不匹配转化为PPO演员更新的OD特定优势 shaping。部署时,该策略仅需一次前向传播。在由考虑随机路径选择的路段传输模型建模的墨尔本干线网络的250个工作日15分钟间隔路段流量轨迹上,对LFPG-RL进行开发与评估。在保留的测试轨迹上,LFPG-RL的均方根误差(RMSE)为4.69,平均绝对百分比误差(MAPE)为20.15%,皮尔逊相关系数为0.995。这些结果表明,与现有方法相比,所提方法是一种更高效、更准确的在线OD需求校准方法。
英文摘要
Dynamic origin-destination (OD) matrix estimation calibrates time-dependent input demand for simulations to reproduce observed link flows. Reinforcement learning is well suited to online estimation because a trained policy estimates demand in a single evaluation, but varying target flows complicate learning from aggregate rewards. We propose reinforcement learning with link-flow propagation guidance (LFPG-RL), combining simulated vehicle propagation records with downstream errors to guide individual OD components. Using 250 weekday link-flow trajectories from a Melbourne arterial network, LFPG-RL achieves a mean test root mean squared error (RMSE) of 7.41 vehicles per 15-minute interval, mean absolute percentage error of 31.24%, and Pearson correlation of 0.987. Comparisons with optimisation, filtering and reinforcement learning without guidance show a 40.1% RMSE reduction over the strongest baseline, gradient descent with propagation guidance. LFPG-RL enables rapid demand updates that keep traffic simulations consistent with observed conditions, supporting the evaluation of traffic management strategies.
发表机构
- Seoul National University(首尔大学)
机构由 AI 辅助整理,请以论文原文为准。