AI 中文总结
该研究提出首个LLM引导的AutoPref框架,将偏好目标分解为成对损失与集合感知加权程序,通过分阶段条件搜索策略实现自动发现,在TSP等四类组合优化问题上性能优于手动设计基线。
AI 中文摘要
组合优化问题(COPs)支撑着诸多现实决策,但其指数级大的搜索空间使得获取高质量解的成本高昂。神经组合优化(NCO)通常通过强化学习(RL)学习快速构造策略,而基于偏好的NCO通过从相对解质量中学习来提高样本效率。然而,现有的偏好目标在手动指定的通用公式中结合了两种不同的设计选择:从每个解对中提取什么学习信号,以及如何相对于采样集对每个对进行加权。我们提出AutoPref,这是第一个用于NCO中自动偏好目标发现的大语言模型(LLM)引导框架。AutoPref将目标分解为成对损失程序(定义学习信号)和集合感知加权程序(确定每个对的相对贡献),它们的组合形成了统一的程序目标空间,其中包含现有偏好目标作为特例。为了使搜索易于处理,我们引入了带有行为门的分阶段条件搜索策略,该策略在短视距训练和评估之前过滤掉不可接受的程序。在旅行商问题(TSP)、带容量约束的车辆路径问题(CVRP)、柔性流水车间调度问题(FFSP)和作业车间调度问题(JSSP)上,AutoPref在所有问题规模上始终优于强大的手动设计基线,证明了NCO的自动目标发现的优势和可扩展性。
英文摘要
Combinatorial optimization problems (COPs) underpin many real-world decisions, but their exponentially large search spaces make high-quality solutions costly to obtain. Neural combinatorial optimization (NCO) learns fast construction policies, typically with reinforcement learning (RL), while preference-based NCO improves sample efficiency by learning from relative solution quality. However, existing preference objectives combine two distinct design choices in manually specified, one-size-fits-all formulations: what learning signal to extract from each solution pair and how to weight each pair relative to the sampled set. We present AutoPref, the first LLM-guided framework for automated preference-objective discovery in NCO. AutoPref factorizes the objective into a pairwise loss program, which defines the learning signal, and a set-aware weighting program, which determines each pair's relative contribution. Their composition forms a unified programmatic objective space containing existing preference objectives as special cases. To make its search tractable, we introduce a staged conditional search strategy with behavioral gates that filter inadmissible programs before short-horizon training and evaluation. Across TSP, CVRP, FFSP, and JSSP, AutoPref consistently outperforms strong hand-designed baselines across problem scales, demonstrating the benefits and scalability of automated objective discovery for NCO.
Comments8pages, 2figures