AI 中文总结
本文提出GPU原生MPC框架CUDA MPC,通过协同设计优化算法等实现低开销,在多机器人基准测试中实时性远超CPU及同类框架,是首个实现10智能体集群无碰撞协调的实时求解器。
AI 中文摘要
模型预测控制(MPC)可实现考虑约束的控制,但它依赖在线优化的特性限制了其在快动态系统、高维模型或长时域场景中的应用。现有GPU实现通常将设备仅作为线性代数加速器,使得优化循环依赖重复的内核启动和高延迟内存传输。本文提出CUDA MPC,一款为CUDA硬件协同设计了优化算法、执行模型和内存架构的GPU原生MPC框架。CUDA MPC将并行时域交替方向乘子法(ADMM)拆分与融合CUDA内核配对,该内核在设备上运行整个迭代求解过程;中间优化变量保留在低延迟的片上共享内存中,局部原子标志协议仅同步相邻时域块,最大程度减少主机干预、内核调度开销和全局内存流量。在六个状态维度和约束密度递增的非线性机器人基准测试中,CUDA MPC在时域比CPU求解器长1至2个数量级的情况下仍保持实时速率:它在0.1秒采样间隔内求解了具有数百步前瞻的基于优化的避障泊车问题,并且是所有被评估求解器中唯一实现实时执行和集中式10智能体集群无碰撞协调的求解器,其中acados和CasADi无法返回可行解,每次求解分别需要3.5秒和4.5秒;与采用相同ADMM拆分的张量框架实现相比,融合内核的速度最高提升965倍。
英文摘要
Model Predictive Control (MPC) delivers constraint-aware control, but its reliance on online optimization limits its use on systems with fast dynamics, high-dimensional models, or long horizons. Existing GPU implementations typically treat the device as a linear-algebra accelerator, leaving the optimization loop dependent on repeated kernel launches and high-latency memory transfers. This paper introduces CUDA MPC, a GPU-native MPC framework that co-designs the optimization algorithm, execution model, and memory architecture for CUDA hardware. CUDA MPC pairs a parallel-in-horizon alternating direction method of multipliers (ADMM) splitting with a fused CUDA kernel that runs the entire iterative solve on the device. Intermediate optimization variables stay in low-latency, on-chip shared memory, and a localized atomic-flag protocol synchronizes only adjacent horizon blocks, minimizing host intervention, kernel-dispatch overhead, and global-memory traffic. Across six nonlinear robotics benchmarks spanning increasing state dimension and constraint density, CUDA MPC sustains real-time rates at horizons one to two orders of magnitude longer than CPU solvers: it solves an optimization-based collision-avoidance parking problem with 100 s of lookahead within a 0.1 s sampling interval, and is the only solver evaluated that achieves both real-time execution and collision-free coordination for a centralized 10-agent swarm, where acados and CasADi return no feasible solution and require 3.5 s and 4.5 s per solve. Against tensor-framework implementations of the same ADMM splitting, the fused kernel is up to $965\times$ faster.