arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LM-GRASP:基于在线模仿学习的组合构造的实例特定语言模型

LM-GRASP: Instance-Specific Language Models for Combinatorial Construction via Online Imitation Learning

Mohand Mezmaz, Grégoire Danoy

arXiv 2607.28135首次发表:更新:

AI 中文总结

本研究提出LM-GRASP框架,将GRASP随机构造阶段转为在线模仿学习任务,用仅解码器Transformer作构造策略,在Taillard PFSP基准上平均比GPU-GRASP提升28.4个制造跨度单位,是手工构造器的实用替代方案。

AI 中文摘要

组合优化领域的机器学习通常依赖于针对固定问题类别的大型离线数据集,通过强化学习训练神经构造器,这会产生高昂的预训练成本,且在训练分布之外的泛化能力较差。我们提出一种替代方案:一种元启发式框架,将GRASP的随机构造阶段重新表述为在线模仿学习任务,针对每个问题实例从头开始训练。局部搜索过程充当专家预言机,而仅解码器的Transformer作为构造策略。与依赖基于局部标量成本的静态近视启发式规则的经典GRASP不同,我们的方法完全是数据驱动的:构造策略从搜索过程中发现的高质量解决方案中产生,无需特定问题的特征工程。我们将其实例化为LM-GRASP,一种遵循迭代学习-推理-改进循环的混合元启发式算法,通过在精英轨迹的动态存档上进行行为克隆在线训练策略,无需外部数据或离线预训练。该流程仅通过局部搜索使用的目标评估器与领域交互。在Taillard PFSP基准(ta51-ta60)上进行评估,该基准因一半最优解未知而成为最具区分度的模块,LM-GRASP平均比GPU-GRASP表现好28.4个制造跨度单位,与GPU加速相对于顺序执行的增益(27.2个单位)相当,尽管存在重叠的标准差。这表明,实例特定的、在线训练的语言模型是手工设计的构造器的有前景的实用替代方案,尤其适用于经典贪婪构造难以处理的场景。

英文摘要

Machine learning for combinatorial optimization typically relies on neural constructors trained via reinforcement learning on large offline datasets for a fixed problem class-incurring high pretraining costs and generalizing poorly outside the training distribution. We propose an alternative: a metaheuristic framework that reformulates the randomized constructive phase of GRASP as an online imitation learning task, trained from scratch on each problem instance. A local search procedure acts as an expert oracle, while a decoder-only Transformer serves as the constructive policy. Unlike classical GRASP, which relies on static, myopic heuristic rules based on localized scalar costs, our approach is fully data-driven: the construction policy emerges from high-quality solutions discovered during the search itself, with no problem-specific feature engineering required. We instantiate this as LM-GRASP, a hybrid metaheuristic following an iterative learn-infer-improve cycle, training the policy online via behavioral cloning on a dynamic archive of elite trajectories-no external data or offline pretraining needed. The pipeline interfaces with the domain solely through the objective evaluator used by local search. Evaluated on the Taillard PFSP benchmark (ta51-ta60), the most discriminating block due to half its optima being unknown, LM-GRASP outperforms GPU-GRASP by 28.4 makespan units on average-comparable to the gain from GPU acceleration over sequential execution (27.2 units), though with overlapping standard deviations. This suggests instance-specific, online-trained language models are a promising, practical alternative to hand-engineered constructors, especially for landscapes resistant to classical greedy construction.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑