arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28988math.OCcs.DC

B³-PWL:面向含SOS2约束的分段线性优化的GPU批处理分支定界算法

B$^3$-PWL: GPU-Batched Branch-and-Bound for Piecewise-Linear Optimization with SOS2 Constraints

  • Johns Hopkins University(约翰斯·霍普金斯大学)
  • University of North Carolina, Chapel Hill(北卡罗来纳大学教堂山分校)
  • City University of Hong Kong(香港城市大学)
  • The Hong Kong Institute of AI for Science(香港人工智能科学研究所)

机构由 AI 辅助整理,请以论文原文为准。

Yilin Guan, Shuqing Luo, Pingzhi Li, Tianlong Chen, Kaidi Xu

中文总结 AI 辅助

本文提出以GPU为中心的B³-PWL框架,通过批处理LP松弛子问题等技术,在PWL优化问题上实现了相对于cuOpt的9.25倍几何平均加速,性能优于多款CPU求解器。

中文摘要 AI 辅助

分段线性(PWL)优化问题出现在众多混合整数规划(MIP)优化应用中,包括投资组合优化、劳动力调度和资源分配。但要将这类问题求解至全局最优仍存在高昂的计算成本,因为分支定界算法需反复求解线性规划(LP)松弛子问题。现有求解器大多以CPU为中心,未能充分利用现代GPU的可扩展性。少数先前的GPU加速分支定界算法要么针对不适合通用PWL优化的神经网络,要么仅加速以CPU为中心的MIP求解器内的辅助子例程(如强分支启发式算法)。为弥合这一差距,我们提出B³-PWL,这是一种以GPU为中心、面向含2型特殊有序集(SOS2)约束的分段线性优化的批处理分支定界框架。我们的方法借助专用的批处理分块稀疏矩阵内核,在GPU上并行求解成批的LP松弛子问题,采用一阶原始对偶求解器。为补充界计算,我们进一步引入统一可行性搜索模块,该模块结合SOS2修复原始启发式算法与批处理可行性泵,以快速获得可行 incumbent 并提升剪枝效率。在包含43个PWL-MIP实例的基准测试中,B³-PWL实现了相对于NVIDIA cuOpt的9.25倍几何平均加速,且在所有测试实例上均获得高质量可行 incumbent。在公开的阀点机组组合基准测试中,它进一步超越了NVIDIA cuOpt以及开源CPU求解器SCIP和HiGHS,证明了一阶LP方法作为GPU加速分支定界核心引擎的潜力。

英文摘要

Piecewise-linear (PWL) optimization problems arise in many mixed-integer programming (MIP) optimization applications, including portfolio optimization, workforce scheduling, and resource allocation. But solving them to global optimality remains computationally expensive because branch-and-bound repeatedly solves LP relaxation subproblems. Existing solvers are largely CPU-centric, leaving the scalability of modern GPUs underutilized. Few prior GPU-accelerated branch-and-bound either targets neural network which is not suitable for general PWL optimization, or accelerates only auxiliary subroutines such as strong branching heuristics within CPU-centric MIP solvers. To bridge this gap, we propose B$^3$-PWL, a GPU-centric batched branch-and-bound framework for piecewise-linear optimization with Special Ordered Set of type 2 (SOS2) constraints. Our method solves batches of LP relaxation subproblems concurrently on the GPU using a first-order primal-dual solver, enabled by a specialized batched block-tiled sparse matrix kernel. To complement bound computation, we further introduce a unified feasibility search module that combines an SOS2 repair primal heuristic with a batched feasibility pump to rapidly obtain feasible incumbents and improve pruning efficiency. On a benchmark of 43 PWL-MIP instances, B$^3$-PWL achieves a 9.25x geometric-mean speedup over NVIDIA cuOpt while reaching high-quality feasible incumbents on every tested instance. On a public valve-point unit-commitment benchmark, it further outperforms NVIDIA cuOpt and the open-source CPU solvers SCIP and HiGHS, demonstrating the potential of first-order LP methods as the central engine of GPU-accelerated branch-and-bound.

补充信息

↑