SPO:通过Stackelberg程序优化发现自适应大邻域搜索算子
SPO: Discovering Adaptive Large Neighborhood Search Operators via Stackelberg Program Optimization
- CDL, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所CDL)
- School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
- Tsinghua University(清华大学)
- University of Chinese Academy of Sciences, Nanjing(中国科学院大学南京学院)
- AiRiA
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
SPO提出基于LLM的Stackelberg程序优化框架,通过状态依赖的破坏-修复程序发现,在TSP和CVRP上超越强基线并泛化到更大实例。
AI中文摘要:
大邻域搜索(LNS)严重依赖于破坏和修复算子,其有效性既取决于对不断演化的LNS状态的适应性,也取决于这两种角色之间的交互。我们引入了Stackelberg程序优化(SPO),一种基于LLM的框架,用于发现自适应的可执行破坏-修复程序。SPO将算子决策条件化于紧凑的LNS状态,使得状态依赖行为能够通过程序发现而涌现,并将破坏-修复发现组织为程序空间上的Stackelberg交互,以反映其不对称依赖关系。角色特定的信用将破坏程序评估为领导者,修复程序评估为条件跟随者响应,指导一个结合了LLM生成器学习与基于种群的进化搜索的耦合优化过程。在旅行商问题和带容量约束的车辆路径问题上的实验表明,SPO在广泛设置中优于强基线,并能泛化到超出发现规模的更大实例和基准集。行为分析进一步展示了发现过程中的状态依赖算子行为和耦合的破坏-修复改进。
英文摘要:
Large neighborhood search (LNS) relies critically on destroy and repair operators, whose effectiveness depends on both adaptation to the evolving LNS state and interaction between the two roles. We introduce Stackelberg Program Optimization (SPO), an LLM-based framework for discovering adaptive executable destroy-repair programs. SPO conditions operator decisions on a compact LNS state, allowing state-dependent behavior to emerge through program discovery, and organizes destroy-repair discovery as a Stackelberg interaction over program space that reflects their asymmetric dependency. Role-specific credits evaluate destroy programs as leaders and repair programs as conditional follower responses, guiding a coupled optimization process that combines LLM generator learning with population-based evolutionary search over programs. Experiments on the traveling salesperson problem and capacitated vehicle routing problem show that SPO outperforms strong baselines across a broad range of settings and generalizes beyond the discovery scale to larger instances and benchmark sets. Behavioral analyses further demonstrate state-dependent operator behavior and coupled destroy-repair improvement during discovery.