发表机构
Tsinghua University; The University of Hong Kong(清华大学; 香港大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对神经组合优化训练中的梯度信号极化与基线冗余问题,提出SSPO方法,通过结构感知相似度加权的留一法基线解决问题,其训练的EFL策略已在京东设施选址系统中实现实际应用。
AI 中文摘要
神经组合优化(NCO)依赖并行解采样进行训练,但现有方法未能充分利用联合采样解组中潜藏的丰富信息。偏好优化方法以单个最优解为锚点,丢弃了所有其他解的细粒度质量与结构信号,这种失败被称为梯度信号极化;而基于均值的基线方法则对所有解均匀加权,导致结构高度相似的解向基线注入冗余信息,使梯度方差维持在较高水平,这种失败被称为基线冗余。我们提出SSPO(结构感知相似度加权偏好优化),它通过基于不相似度的留一法基线对所有B个采样解进行联合评分:结构差异显著的解会获得更高权重,通过单一机制同时解决上述两种问题。该基线使用零参数、问题自适应的解嵌入,这些嵌入由编码器现有的节点表示构建而成。在TSP、EFL和JSP基准上的实验显示,与现有的最优锚点和均匀加权基线相比,SSPO取得了持续的性能提升;在TSP和EFL上与均匀RLOO的直接对比证实,结构感知加权是性能提升的主要驱动因素。经SSPO训练的EFL策略已部署在京东(JD.com)的生产设施选址系统中,验证了其大规模应用的实际可行性。
英文摘要
Neural combinatorial optimization (NCO) relies on parallel solution sampling for training, yet existing methods fail to fully exploit the rich information latent in a co-sampled solution group. Preference-optimization methods anchor on the single best solution and discard fine-grained quality and structural signal from all other peers-a failure we term gradient signal polarization. Mean-based baselines instead weight peers uniformly, so structurally near-identical peers flood the baseline with redundant information and keep gradient variance high-a failure we term baseline redundancy. We propose SSPO (Structure-Aware Similarity-Weighted Preference Optimization), which scores all $B$ sampled solutions jointly through a dissimilarity-weighted leave-one-out baseline: structurally distinct peers receive higher weight, resolving both failures in a single mechanism. The baseline uses zero-parameter, problem-adaptive solution embeddings built from the encoder's existing node representations. Experiments on TSP, EFL, and JSP benchmarks show consistent gains over prior best-anchor and uniform-weight baselines. A direct comparison against uniform RLOO on TSP and EFL confirms that structure-aware weighting is the primary driver of improvement. The SSPO-trained EFL policy has been deployed in a production facility-location system at JD$\mathord{.}$com, confirming practical viability at scale.