HyCO:一种用于组合优化的混合神经求解器
HyCO: A Hybrid Neural Solver for Combinatorial Optimization
浏览论文内容
中文总结 AI 辅助
针对神经组合优化中RL求解器与扩散模型求解器各自的遗憾缺陷,提出混合求解器HyCO,通过RL构建前缀并自适应切换至扩散模型,理论证明降低遗憾并实验验证改进。
中文摘要 AI 辅助
在优化遗憾视角下,用于神经组合优化的序列强化学习(RL)求解器和全局扩散模型(DM)求解器表现出互补的失败模式。前者在早期构建阶段享有较小的边际遗憾,但遭受随视野累积的复合误差,遗憾呈超线性增长;后者避免了视野复合误差,但相对于剩余未解子空间的维度产生线性或次线性遗憾。我们提出了用于组合优化的混合神经求解器(HyCO),一种混合推理算法,它使用RL求解器构建解的前缀,并自适应地切换到条件扩散模型以完成剩余决策。为了刻画这种混合为何有帮助、何时触发切换以及如何在实践中实现,我们首先开发了一个统一的误差缩放理论框架,并证明在显式误差缩放假设下:i)混合结构比任一单独骨干实现严格更低的期望遗憾,ii)存在一个唯一的最优触发步骤以最小化混合遗憾。然后,我们设计了一个轻量级自适应触发器,结合策略熵和RL-DM分歧来检测轨迹级别的机制转变信号作为实际代理,因为最优触发步骤是在期望遗憾级别定义的,无法在单个轨迹上直接计算。在多种基准上的实验结果表明,HyCO在两个骨干上均实现了一致的改进,并支持自适应触发的经验有效性。
英文摘要
Sequential reinforcement learning (RL) solvers and global diffusion model (DM) solvers for neural combinatorial optimization exhibit complementary failure modes under an optimization-regret view. The former enjoys small marginal regret in the early construction stage, but suffers from horizon-wise compounding errors with super-linear regret growth; the latter avoids horizon compounding but incurs linear or sublinear regret w.r.t. the dimension of the remaining unsolved subspace. We propose Hybrid Neural Solver for Combinatorial Optimization (HyCO), a hybrid inference algorithm that constructs a solution prefix with an RL solver and adaptively switches to a conditional DM to complete the remaining decisions. To characterize why such hybridization helps, when to trigger the handover, and how to realize it in practice, we first develop a unified error-scaling theoretical framework and prove that, under explicit error-scaling assumptions, i) the hybrid structure achieves strictly lower expected regret than either backbone alone, and ii) there exists a unique optimal trigger step that minimizes the hybrid regret. We then design a lightweight adaptive trigger that combines policy entropy and RL-DM disagreement to detect trajectory-level signals of the regime shift as a practical proxy, since the optimal trigger step is defined at the expected-regret level and is not directly computable on individual trajectories. Experimental results on diverse benchmarks demonstrate that HyCO achieves consistent improvements over both backbones and support the empirical effectiveness of adaptive triggering.
发表机构
- College of William & Mary(威廉玛丽学院)
机构由 AI 辅助整理,请以论文原文为准。