发表机构
Southern University of Science and Technology; Guangdong Provincial Key Laboratory of Fully Actuated System Control Theory and Technology; Pengcheng Laboratory(南方科技大学; 广东省全驱系统控制理论与技术重点实验室; 鹏城实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MiLoop提出纯强化学习的构造性框架,通过选择性记忆传播实现动态嵌入,无需解标签或剪枝,在100至千万节点实例上展现强泛化能力。
AI 中文摘要
构造性神经组合优化(NCO)已成为一种有前景的范式,它学习逐步构造组合优化问题(COPs)的解,从而减少对手工规则的依赖并实现快速推理。尽管许多采用动态嵌入的方法具有良好的泛化能力,但它们通常每一步都使用深层注意力堆栈从头重建子问题表示。该类别中许多高性能方法依赖解标签或伪标签进行高效训练,或在强化学习(RL)期间进行激进的搜索空间剪枝。为解决这些局限性,我们提出了记忆在环(MiLoop),一种纯基于RL的构造性框架,利用rollout所需的多步计算进行选择性记忆传播。每次rollout为学习提供解质量反馈,同时传播历史表示,从而使浅层策略无需外部解标签或训练时搜索空间剪枝即可学习有效的动态嵌入。具体而言,MiLoop在注意力层之前将当前嵌入与历史记忆融合,并在之后应用自适应门控更新。更新后的表示同时支持当前决策和逐步重用。在四个COP上的大量实验表明,MiLoop在从100到1000万个节点的实例上持续产生高质量解,突显了其强大的泛化能力。
英文摘要
Constructive neural combinatorial optimization (NCO) has emerged as a promising paradigm that learns to construct solutions to combinatorial optimization problems (COPs) step by step, which reduces reliance on handcrafted rules and enables fast inference. While many methods with dynamic embeddings generalize well, they typically rebuild subproblem representations from scratch at each step using deep attention stacks. Many high-performing methods in this category rely on solution labels or pseudo-labels for efficient training, or on aggressive search space pruning during reinforcement learning (RL). To address these limitations, we propose Memory-in-the-Loop (MiLoop), a purely RL-based constructive framework that leverages the multi-step computation already required by a rollout for selective memory propagation. Each rollout provides solution-quality feedback for learning while propagating historical representations, thereby enabling a shallow policy to learn effective dynamic embeddings without external solution labels or training-time search-space pruning. Specifically, MiLoop fuses current embeddings with historical memory before the attention layers and applies adaptive gated updates afterward. The updated representations support both current decisions and stepwise reuse. Extensive experiments across four COPs demonstrate that MiLoop consistently produces high-quality solutions on instances ranging from 100 to 10 million nodes, highlighting its strong generalization ability.