arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GPU集群上极端规模下的分布式线性规划

Distributed Linear Programming on GPU Clusters at Extreme Scale

Arnaud Deza, Santanu Dey, Pascal Van Hentenryck

arXiv 2609.09108首次发表:更新:

AI 中文总结

SHARDLP是一种分布式GPU线性规划求解器,通过全流程分区和通信优化,在极端规模下显著提升求解效率,并在基准测试中大幅超越现有CPU方法。

AI 中文摘要

大型线性规划问题可能超出单个计算节点的内存容量。尽管一阶方法用适合GPU的矩阵-向量乘积取代了稀疏分解,但求解器的其他阶段可能重新引入单节点内存限制。我们提出SHARDLP,一种分布式GPU线性规划求解器,它从分片输入到解输出,始终保持矩阵和原始-对偶状态的分区。在Google PDLP基准测试上,SHARDLP在十一个实例中的九个达到了已发表的标准,而已发表的CPU PDLP研究为八个。在最大的基准测试上,八个H200 GPU在9.9分钟内解决了一个具有11.85亿变量和63.38亿非零元的线性规划问题;已发表的CPU实验报告在不同硬件上耗时21.06小时。超出此基准,单独验证的多节点求解达到高达136.04亿变量和408.07亿非零元,而经过验证的执行跨越多达29个计算节点上的76个GPU。对于列分区求解,支持感知通信跳过那些对某行不存储系数的GPU;在一个具有27.6亿非零元的线性规划上,它将建模通信减少了92.97%,并将求解时间提高了1.27倍至1.52倍。

英文摘要

Large linear programs can exceed the memory of a single compute node. Although first-order methods replace sparse factorizations with GPU-suited matrix-vector products, other solver phases can reintroduce a single-node memory limit. We present SHARDLP, a distributed GPU LP solver that keeps the matrix and primal-dual state partitioned from sharded input through solution output. On the Google PDLP benchmark, SHARDLP reaches the published criterion on nine of eleven instances, compared with eight in the published CPU PDLP study. On the largest benchmark, eight H200 GPUs solve a 1.185-billion-variable, 6.338-billion-nonzero LP in 9.9 minutes; the published CPU experiment reports 21.06 hours on different hardware. Beyond this benchmark, separately checked multi-node solves reach up to 13.604 billion variables and 40.807 billion nonzeros, while validated executions span up to 76 GPUs across 29 compute nodes. For column-partitioned solves, support-aware communication skips GPUs that store no coefficients for a row; on an LP with 2.76 billion nonzeros, it cuts modelled communication by 92.97% and improves solver time by 1.27x-1.52x

Comments18 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑