arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24859cs.LG

通过局部增强偏好与表示解耦改进跨问题车辆路径规划

Improving Cross-Problem Vehicle Routing with Locally Augmented Preferences and Representation Disentanglement

Arthur Corrêa, Paulo Nascimento, Samuel Moniz

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对多任务VRP求解器的训练与架构缺陷,提出POLAR算法与PLE编码器,在16个分布内变体上平均差距降21.3%,27个未见变体优于现有方法,提升了跨问题泛化性能。

中文摘要 AI 辅助

多任务车辆路径规划问题(VRP)求解器旨在用单个统一模型处理多种VRP变体,避免为每种变体单独训练模型。尽管近期取得进展,现有方法仍存在两方面局限:训练层面,强化学习存在奖励尺度差异,且随着策略优化优势信号逐渐收缩;偏好优化在采样路径近乎完全相同时会陷入停滞,根本上受限于策略自身生成解的质量,导致两种范式在训练过程中监督信号均较弱。架构层面,现有完全共享的编码器会将异构变体中依赖约束的表示纠缠,限制了泛化能力。我们通过两项与模型无关的贡献解决这些缺陷:其一,提出局部增强优化偏好(POLAR)算法,这是一种新型训练算法,在形成偏好对前,对最优解码路径应用局部搜索优化步骤,从而产生更具信息量的成对边际;其二,渐进式分层提取(PLE)编码器,通过门控机制使每个编码器层经过一个共享专家和一组特定任务专家,逐步将通用路径结构与依赖约束的编码分离。通过在多种VRP变体上开展大量实验,我们表明POLAR与PLE结合提升了神经多任务求解器的当前最优水平:在16个分布内变体上,相比最强已发表基线,我们将与参考解的平均差距降低了21.3%;在32个未见变体中,有27个的性能优于现有神经方法。消融研究证实了每项贡献的有效性,表明二者均能在多种骨干模型架构上提升跨问题泛化能力。

英文摘要

Multi-task vehicle routing problem (VRP) solvers seek to handle multiple VRP variants within a single unified model, avoiding the need to train a separate model for every variant. In spite of recent progress, current approaches remain limited on two fronts. On the training side, reinforcement learning suffers from reward-scale disparities and shrinking advantage signals as policies improve, whereas preference optimization stagnates once sampled tours become near-identical and thus fundamentally limited by the quality of the policy's own generated solutions, leaving both paradigms with weak supervision as training progresses. On the architecture side, existing fully shared encoders entangle constraint-dependent representations across heterogeneous variants, which limits generalization. We address these gaps with two model-agnostic contributions. First, we propose Preference Optimization with Locally Augmented Refinement (POLAR), a novel training algorithm that applies a local search refinement pass to the best decoded tour before forming preference pairs, yielding much more informative pairwise margins. Second, a Progressive Layered Extraction (PLE) encoder routes each encoder layer through one shared expert and a set of task-specific experts via a gating mechanism, progressively separating common routing structure from constraint-specific encodings. Through extensive experiments on various VRP variants, we show that POLAR and PLE together elevate the current state-of-the-art among neural multi-task solvers. We reduce the average gap to reference solutions by 21.3% relative to the strongest published baseline on 16 in-distribution variants, and outperform prior neural methods on 27 out of 32 unseen variants. Ablation studies confirm the efficacy of each contribution, showing that both improve cross-problem generalization across multiple backbone model architectures.

发表机构

  • University of Coimbra(科英布拉大学)

机构由 AI 辅助整理,请以论文原文为准。

↑