WarpMPC:通过展开的$LDL^\top$分解和交替方向乘子法在GPU上进行大批量模型预测控制
WarpMPC: Large-Batch MPC on GPU via ADMM with Unrolled $LDL^\top$ Factorization
- Institute for Data Science in Mechanical Engineering (DSME), RWTH Aachen University(亚琛工业大学机械工程数据科学研究所)
- Department of Mechanical Engineering, Massachusetts Institute of Technology(麻省理工学院机械工程系)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究在GPU上求解大批量相同结构的顺序二次规划迭代以最大化吞吐量的问题,核心方法是基于问题实例稀疏性提出展开稀疏线性分解等优化,主要贡献是在多机器人基准测试中大幅提升吞吐量并展示实际应用价值。
AI中文摘要:
本文介绍了在求解大批量(10000至超过100000)具有相同结构的顺序二次规划(SQP)迭代时,为最大化GPU吞吐量而进行的数值优化。这些优化在用于模型预测控制(MPC)的WarpMPC工具箱中实现。基于一批中所有MPC问题实例在时间、成本和约束上具有相同稀疏性的见解,提出展开稀疏线性分解并求解,避免内存访问瓶颈和计算浪费,加速灵敏度计算。在非线性小车、四旋翼和人形机器人基准测试中实现每秒8000至250000次SQP迭代的吞吐量,优于基线3至25倍。通过合成数据集并在4分钟内训练MPC的神经网络近似,展示了其实际用途。
英文摘要:
This paper introduces numerical optimizations for maximizing throughput on GPU when solving large batches (10,000 to over 100,000) of sequential quadratic programming (SQP) iterations, where all problems have the same structure. The optimizations are implemented in a toolbox WarpMPC for model-predictive control (MPC) in JAX and Warp. Based on the insight that all MPC problem instances in a batch share the same sparsity in time, cost, and constraints, we propose unrolling sparse linear factorizations and solves, which dominate alternating direction method of multipliers (ADMM) solver runtime. We avoid memory access bottlenecks and wasting computations via optimized memory layout, padding-reducing segmentation of the unrolled factorization, and dependency level scheduled backsolves, additionally accelerating sensitivity computation. We achieve throughputs of 8,000 to 250,000 SQP iterations per second on nonlinear cartpole, quadrotor, and humanoid robot benchmarks, outperforming baselines by 3$\times$ to 25$\times$. We illustrate practical usefulness by synthesizing a dataset and training a neural network approximation of an MPC in under 4 minutes that stabilizes a nano quadrotor in hardware experiments.