发表机构
Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出GALLOP,用强化学习联合学习PDHG算法的连续参数与离散重启决策,在六个LP族上减少1.9-5.6倍迭代次数,最高16倍加速,且策略可泛化至更大规模问题。
AI 中文摘要
原对偶混合梯度(PDHG)方法利用适合GPU的矩阵-向量乘积和投影来求解大规模线性规划(LP),但其实际性能依赖于算法参数、加速和重启的协调。我们提出了GALLOP,该方法使用强化学习来联合学习连续算法参数和离散重启决策,而无需对求解器进行微分。其广义加速PDHG更新结合了独立的原始和对偶外推、历史修正以及具有独立可调系数的重启锚定。我们使用分组近端策略优化目标训练一个维度无关的反馈策略,该目标分别对不同控制组裁剪似然比,并在重启转换时排除非活动的加速控制。我们在六个LP族和一个公共物品放置基准上评估了GALLOP。在六个族的主要评估设置中,GALLOP将迭代次数减少了1.9至5.6倍,并在算法墙钟时间上实现了相对于MPAX高达16.0倍的加速。每个族训练一个策略,学习到的策略无需重新训练即可泛化到比最大训练实例大3倍至400倍的族内LP,包括具有1024万个变量的运输LP。
英文摘要
Primal-dual hybrid gradient (PDHG) methods solve large-scale linear programs (LPs) using GPU-friendly matrix-vector products and projections, but their practical performance depends on coordinating algorithm parameters, acceleration, and restarts. We introduce GALLOP, which uses reinforcement learning to jointly learn continuous algorithm parameters and discrete restart decisions without differentiating through the solver. Its generalized accelerated PDHG update combines separate primal and dual extrapolation, history corrections, and restart anchoring with independently adjustable coefficients. We train a dimension-agnostic feedback policy using a groupwise proximal policy optimization objective that clips likelihood ratios separately for different control groups and excludes inactive acceleration controls on restart transitions. We evaluate GALLOP on six LP families and a public item-placement benchmark. On the main evaluation settings across the six families, GALLOP reduces iteration counts by factors of $1.9$-$5.6$ and achieves up to a $16.0\times$ speedup in algorithm wall-clock time over MPAX. With one policy trained per family, the learned policies generalize without retraining to within-family LPs $3\times$-$400\times$ larger than the largest training instances, including Transport LPs with $10.24$ million variables.
Comments35 pages, 4 figures