AI 中文总结
研究针对大规模半定规划提出CARDAL求解器,基于秩自适应分解和增广拉格朗日方法,通过特定操作和问题选择实现优化,分析了全局最优性,在基准测试和多领域实验中展现强鲁棒性与加速效果。
AI 中文摘要
我们提出了CARDAL(曲率感知秩自适应分布式增广拉格朗日),一种基于秩自适应Burer-Monteiro分解和增广拉格朗日方法的大规模半定规划(SDP)分布式多GPU求解器。在固定秩时,带负曲率校正的无矩阵L-BFGS方法针对分解后的增广拉格朗日的近似欧几里得二阶驻点。反向乘子移位在秩扩展后将负对偶松弛方向转变为精确负曲率,一个小的联合秩提升问题选择批量低秩校正。一个经过验证的松弛下界提供后验近似KKT证书。我们的分析为接近Barvinok-Pataki尺度的每个块秩处的PSD锥的异构乘积建立了通用全局最优性保证,以及在逐块成本平滑下的有限精度对应。为了可扩展执行,CARDAL在约束x秩x锥设备网格上分布约束行因子列和PSD块。在Mittelmann基准测试中,在统一精度标准下,CARDAL比现有的低秩GPU方法表现出更强的鲁棒性。在机器人技术、电子结构和最大割的大规模SDP松弛实验中,展示了三个分布轴的互补缩放机制,在四个H100 GPU上观察到高达4倍的壁钟加速。
英文摘要
We present CARDAL (Curvature-Aware Rank-Adaptive Distributed Augmented Lagrangian), a distributed multi-GPU solver for large-scale semidefinite programs (SDPs) based on a rank-adaptive Burer-Monteiro factorization and an augmented Lagrangian method. At fixed ranks, a matrix-free L-BFGS method with negative-curvature corrections targets an approximate Euclidean second-order stationary point of the factored augmented Lagrangian. A reverse multiplier shift turns a negative dual-slack direction into exact negative curvature after rank expansion, and a small joint rank-lift problem selects a batched low-rank correction. A verified slack lower bound provides an a posteriori approximate KKT certificate. Our analysis establishes generic global-optimality guarantees for heterogeneous products of PSD cones at per-block ranks near the Barvinok-Pataki scale, together with a finite-accuracy counterpart under blockwise cost smoothing. For scalable execution, CARDAL distributes constraint rows, factor columns, and PSD blocks over a Constraint x Rank x Cone device mesh. The primal residual, gradient, Hessian-vector products, and slack matrix-vector products are evaluated using device-local operations and axis-wise collectives. On the Mittelmann benchmark, CARDAL exhibits stronger robustness than existing low-rank GPU approaches under a uniform accuracy standard. Experiments on large-scale SDP relaxations from robotics, electronic structure, and Max-Cut demonstrate the complementary scaling regimes of the three distribution axes, with observed wall-clock speedups of up to 4x on four H100 GPUs.
Comments25 pages, 5 tables