带在线请求的动态多车场车辆路径问题:事件驱动Transformer深度强化学习与滚动时域基准测试
Dynamic Multi-Depot Vehicle Routing with Online Requests: Event-Driven Transformer--DRL and Rolling-Horizon Benchmarking
- Louisiana State University(路易斯安那州立大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对带在线请求的动态多车场车辆路径问题,提出事件驱动的Transformer深度强化学习框架,经多场景基准测试,该学习策略在部分指标上弱于最强启发式算法,无单一方法在所有维度最优。
AI中文摘要:
本文针对请求逐步揭示、车辆状态不断变化的动态多车场车辆路径问题,提出了一种事件驱动的学习与基准测试框架。通过行为克隆和近端策略优化(PPO)训练了掩码多层感知机(Masked MLP)和Transformer策略;确定性可行性掩码可防止车辆与请求的无效分配,而固定前缀/灵活后缀的路径承诺则可保护已完成、活跃及近期决策,并分别衡量车辆重新分配与重新排序情况。将学习到的策略与动态插入启发式算法、限时滚动时域优化进行对比:在20个场景的策略基准测试中,所有方法均完成了所有请求且无无效动作,但最近可行方法的平均目标值最低,在路径质量、等待时间、稳定性、总完成时间(makespan)和运行时间方面均优于学习到的策略;在5次独立训练运行中,PPO对MLP的平均影响很小,对Transformer平均有所改进,但种子变异性更大;在通用协议下,最近可行方法的综合目标值和路径中断程度最低,而滚动时域方法的等待时间和总完成时间最低,但计算成本显著更高;学习到的策略保持了毫秒级决策速度,无需重新训练即可迁移到最多含80个请求的实例,但未优于最强启发式算法;在路径效率、服务响应性、稳定性和在线计算方面,没有单一方法表现最优。
英文摘要:
This paper presents an event-driven learning and benchmarking framework for the Dynamic Multi-Depot Vehicle Routing Problem with progressively revealed requests and evolving vehicle states. Masked MLP and Transformer policies are trained through behavior cloning and proximal policy optimization. Deterministic feasibility masking prevents invalid vehicle--request assignments, while fixed-prefix/flexible-suffix route commitments protect completed, active, and near-term decisions and separately measure vehicle reassignment and resequencing. The learned policies are compared with dynamic insertion heuristics and time-limited rolling-horizon optimization. In a 20-scenario policy benchmark, all methods completed every request without invalid actions, but nearest feasible achieved the lowest mean objective and outperformed the learned policies in routing quality, waiting time, stability, makespan, and runtime. Across five independent training runs, PPO had little average effect on the MLP and improved the Transformer on average, although with greater seed variability. Under the common protocol, nearest feasible achieved the lowest combined objective and route disruption, whereas rolling horizon achieved the lowest waiting times and makespan at substantially higher computational cost. The learned policies retained millisecond-level decisions and transferred to instances with up to 80 requests without retraining, but did not outperform the strongest heuristic. No single method was best across routing efficiency, service responsiveness, stability, and online computation.