arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在 GPU 时代重新审视动态优化的同步方法

Revisiting Simultaneous Methods for Dynamic Optimization in the GPU Era

Joseph W. Choi, Sungho Shin

arXiv 2607.12201首次发表:更新:

AI 中文总结

该研究重新审视动态优化中解决 DAE 约束优化问题的顺序与同步方法,在 GPU 计算环境下评估同步方法性能。采用基于正交配置的同步方法用于系统生物学参数估计基准测试,结果显示其在 GPU 上有优势,与顺序基线相比加速比可达 5.4 倍。

AI 中文摘要

我们重新审视动态优化中的经典话题:解决 DAE 约束优化问题的顺序方法与同步方法,特别关注图形处理单元(GPU)计算如何改变它们的有效性。顺序方法通过在微分代数方程(DAE)求解器级别进行自适应时间步长提供关键优势,对处理刚性系统特别有效,但长时间模拟仍是计算瓶颈。同步方法适合并行计算,通过利用离散 DAE 系统在函数评估级别的高度重复结构和稀疏线性代数例程实现消除树级并行。本文在 GPU 计算环境中重新审视同步方法的能力,并与顺序方法基线评估其性能。我们采用基于正交配置的同步方法应用于系统生物学的参数估计基准测试,评估 GPU 和 CPU 求解器。结果表明,同步方法虽可靠性较低,但在两种方法都成功求解的最大实例中,与顺序基线相比加速比高达 5.4 倍,且在 GPU 上相对于问题大小的优势比在 CPU 上更明显。

英文摘要

We revisit the classical topic in dynamic optimization: sequential vs simultaneous methods for solving DAE-constrained optimization problems, with a particular focus on how graphics processing unit (GPU) computing changes their effectiveness. Sequential methods offer key advantages through adaptive time stepping at the differential-algebraic equation (DAE) solver level, which is especially effective for handling stiff systems. However, long-time-horizon simulations remain a computational bottleneck, as time integration is inherently sequential and limits parallelization within the optimization algorithm. In contrast, simultaneous approaches are well-suited for parallel computing. They address these limitations by exploiting the highly repetitive structure of discretized DAE systems at the function evaluation level and leveraging sparse linear algebra routines that enable elimination tree-level parallelism. Although simultaneous methods typically lack adaptive time stepping, this limitation can often be mitigated by choosing a sufficiently fine initial mesh or iteratively adjusting mesh coarseness in an outer loop. In this work, we revisit the capabilities of the simultaneous approach in a GPU computing environment and assess its performance against a sequential method baseline. We employ a simultaneous approach based on orthogonal collocation within an open-source modeling framework and apply it to parameter estimation benchmarks from systems biology. We evaluate both GPU and CPU solvers on the simultaneous formulation. Our results show that, although less reliable, the simultaneous approach achieves up to 5.4x speedup compared to the sequential baseline among the largest instances where both methods solve successfully. The advantage of the simultaneous method with respect to problem size is more pronounced on GPUs than on CPUs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑