AI 中文总结
针对加拿大林业行业原木卡车路径规划和调度在干扰后的实时重新优化问题,提出将恢复制定为参数化的邻域受限混合整数线性规划,并用强化学习策略自适应选择参数值以高效导航稳定性 - 成本空间,能产生更优可行恢复计划。
AI 中文摘要
我们考虑加拿大林业行业在意外干扰后原木卡车路径规划和调度的实时重新优化。道路封闭、车辆故障、行程延误和需求波动使预先制定的战术计划无效,需要恢复决策以同时恢复可行性并限制与原计划的偏差,这两个目标在紧迫时间限制下内在矛盾。我们将恢复制定为一系列由偏差界限$k$参数化的邻域受限混合整数线性规划,$k$是恢复计划与基线计划之间的$\ell_1$距离。我们表明可以精确计算最小可行值$k_{\min}$,将搜索锚定在最稳定的极端。我们寻求跨越稳定性 - 成本空间的非支配恢复计划的帕累托前沿,为调度员提供结构化的操作选项集。为了有效导航该空间,我们提出一种强化学习策略,用REINFORCE策略梯度算法训练,它根据求解器反馈自适应选择连续的$k$值,将计算精力集中在有效区域,并在不太可能进一步改进时终止探索。在从加拿大林业合作伙伴的历史数据得出的每周实例上进行评估,在涵盖所有考虑事件类别的干扰场景下,该方法比固定$k$网格和二分搜索基线用更少次求解器调用恢复更丰富的帕累托前沿,并在所有配置和干扰类型的可操作接受的重新优化时间内产生可行的恢复计划。
英文摘要
We consider the real-time reoptimisation of log-truck routing and scheduling in the Canadian forestry industry following unforeseen disruptions. Road closures, vehicle breakdowns, travel delays, and demand fluctuations invalidate pre-established tactical plans and call for recovery decisions that simultaneously restore feasibility and limit deviation from the original schedule, two objectives inherently in tension under tight time constraints. We formulate recovery as a sequence of neighbourhood-restricted mixed-integer linear programs parameterised by a deviation bound $k$, the $\ell_1$ distance between the recovered and baseline plans. We show that the minimum feasible value $k_{\min}$ can be computed exactly, anchoring the search at its most stable extreme. Rather than targeting a single $k$, we seek a Pareto front of non-dominated recovery plans spanning the stability--cost space, giving dispatchers a structured set of operational options. To navigate this space efficiently, we propose a reinforcement learning policy, trained with the REINFORCE policy-gradient algorithm, that adaptively selects successive values of $k$ from solver feedback, concentrating computational effort in productive regions and terminating exploration when further improvement is unlikely. Evaluated on weekly instances derived from historical data of a Canadian forestry partner, under disruption scenarios covering all event categories considered, the approach recovers richer Pareto fronts with fewer solver calls than fixed-$k$ grid and dichotomic search baselines, and produces feasible recovery plans within operationally acceptable reoptimization times across all configurations and disruption types.