arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种用于具有随机需求和外包的车辆路径问题的深度强化学习算法

A Deep Reinforcement Learning Algorithm for the Vehicle Routing Problem with Stochastic Demands and Outsourcing

Mohsen Dastpak, Fausto Errico, Ola Jabali

arXiv 2607.16875首次发表:更新:

发表机构

Department de génie de la construction, École de technologie supérieure(土木工程系,超技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究具有随机需求和外包选项的车辆路径问题,提出迭代两级方法,第一级划分客户子集,第二级用图注意力网络表示状态的深度Q网络求解VRP-SD成本,实验表明该方法能大幅降低路径成本,快速生成高质量决策。

AI 中文摘要

我们引入了具有随机需求和外包选项的车辆路径问题(VRP-SDO),物流服务提供商将客户请求分为外包给公共承运人及由其固定车队负责的客户。后者引发了带随机需求的车辆路径问题(VRP-SD)并动态求解。需求在访问时揭示,剩余需求可由其他车辆或在仓库补货后服务。超常规班次工作产生加班成本,单位外包成本随预期外包需求降低。目标是最小化预期行程、加班和外包成本。我们提出一种迭代两级方法,第一级将客户分为负责和外包子集,第二级估计预期VRP-SD路径成本。为避免每次迭代从头求解,学习离线路径策略以快速估计成本。迭代局部搜索确定第一级划分。将第二级公式化为马尔可夫决策过程并用深度Q网络求解,其状态由图注意力网络表示,通过与作业车辆的相关性聚合客户和车辆信息。在具有可变客户数量和位置的实例上离线训练,该策略适用于任何每日客户情况;在线微调改善成本近似。实验表明,我们的策略相对于最先进方法将路径成本降低19.6%,比经典启发式方法至少降低29.6%。我们的整体算法比无注意力表示的版本平均节省13.7%,并在几分钟内生成高质量决策,而无离线训练估计器的基准测试需要超过一小时。

英文摘要

We introduce the vehicle routing problem with stochastic demands and outsourcing options (VRP-SDO), in which a logistics service provider partitions customer requests into customers outsourced to a common carrier and customers committed to its fixed fleet. The latter induces a vehicle routing problem with stochastic demands (VRP-SD), solved dynamically. Demands are revealed upon visit; residual demand may be served by other vehicles or after restocking at the depot. Work beyond the regular shift incurs overtime costs, and the unit outsourcing cost decreases with the expected outsourced demand. The objective is to minimize expected travel, overtime, and outsourcing costs. We propose an iterative two-level methodology whose first level partitions customers into committed and outsourced subsets, while the second level estimates the expected VRP-SD routing cost. To avoid solving this problem from scratch at every iteration, we learn an offline routing policy that estimates costs almost instantly for any committed subset. An iterated local search establishes the first-level partitions. We formulate the second level as a Markov decision process and solve it with a deep Q-network whose state is represented by a graph attention network aggregating customer and vehicle information by relevance to the acting vehicle. Trained offline on instances with variable customer cardinality and locations, the policy applies to any daily customer realization; online fine-tuning improves the cost approximation. Experiments show that our policy reduces routing costs by 19.6% relative to a state-of-the-art method and by at least 29.6% over classical heuristics. Our overall algorithm saves 13.7% on average over the version without the attention-based representation and generates high-quality decisions within minutes, whereas benchmarks without an offline-trained estimator require over an hour.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑