飓风中断下的分货多智能体强化学习联运货运路径规划
Per-Shipment Multi-Agent Reinforcement Learning for Intermodal Freight Routing Under Hurricane Disruption
浏览论文内容
中文总结 AI 辅助
针对飓风破坏下的联运货运路径规划问题,研究将其建模为Dec-POMDP,训练IPPO并与启发式方法对比,发现IPPO在吞吐量和交付率上占优,在需求激增时优势更显著,MAPPO则存在队列不匹配的局限性。
中文摘要 AI 辅助
联运货运网络面临着气候极端事件带来的日益增长的中断风险,这些事件会同时破坏多个运输通道。为解决这一问题,我们将货运路径规划建模为具有分货行动粒度的 Dec-POMDP(分散式部分可观测马尔可夫决策过程),并采用集中式训练与分散式执行框架训练独立近端策略优化算法(IPPO),在飓风中断场景下的15枢纽网络上,与两个拥有特权状态信息的启发式基线进行对比。在30次匹配的回合中,没有任何单一策略占据绝对优势:IPPO实现了最高的吞吐量(+12.7%)和交付率,而考虑容量的启发式方法在弹性指数(RI)和延迟方面表现更优。在需求激增(容量比为2.9:1)的情况下,IPPO的RI优势扩大至+6.4%,表明在容量稀缺时,学习得到的路径规划最具价值。多智能体近端策略优化算法(MAPPO)变体在训练-评估队列不匹配时出现性能崩溃(RI=0.811);重新训练可将RI恢复至1.018,但IPPO在吞吐量上仍占据优势,这表明在分货调度下,集中式评论者仍存在残余局限性。
英文摘要
Intermodal freight networks face growing disruption risk from climate extremes that degrade multiple corridors simultaneously. To address this, we formulate freight routing as a Dec-POMDP with per-shipment action granularity and train Independent PPO (IPPO) under Centralized Training with Decentralized Execution, comparing against two heuristic baselines with privileged state access on a 15-hub network under hurricane disruption. Across 30 matched episodes, no single policy dominates: IPPO achieves the highest throughput ($+12.7\%$) and delivery rate while a capacity-aware heuristic leads on Resilience Index (RI) and delay. Under demand surge (2.9:1 capacity ratio), IPPO's RI advantage grows to $+6.4\%$, suggesting learned routing is most valuable when capacity is scarce. A Multi-Agent PPO (MAPPO) variant collapses under train-eval queue mismatch ($\mathrm{RI} = 0.811$); retraining recovers RI to $1.018$ but IPPO still leads on throughput, pointing to residual limitations in centralized critics under per-shipment dispatch.