arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PETSc/TAO中一种可移植且通用的有限内存BFGS实现

A Portable and Versatile Limited-Memory BFGS Implementation in PETSc/TAO

Hansol Suh, Tobin Isaac, Alp Dener, Todd Munson, Hong Zhang, Richard Tran Mills

arXiv 2608.00196首次发表:更新:

AI 中文总结

本文在PETSc/TAO中实现了兼具递归与紧凑密集L-BFGS优势的中间密集BFGS形式,并在Polaris和Frontier机器上完成了单节点GPU与CPU性能测试。

AI 中文摘要

有限内存BFGS(L-BFGS)海森更新方案是许多拟牛顿优化算法的核心。当存在m个历史步时,实现L-BFGS最常用的方法是在求解线性系统时采用2m个顺序秩1更新。当同步延迟显著时,该方法的性能会下降,且当向量无法放入缓存时,其较差的时间局部性会增加内存流量。L-BFGS的紧凑密集表示法可实现最小的同步延迟和更好的时间局部性,但需要额外遍历基向量,且当B₀矩阵在可变度量方法中发生变化时,必须重新计算额外的基。在便携式科学计算可扩展工具包(PETSc)和高级优化工具包(TAO,即PETSc/TAO)中,我们实现了一种BFGS的中间密集形式,保留了递归方法和紧凑密集方法的大部分优良特性。我们在美国能源部的Polaris和Frontier机器上对这些实现进行了单节点性能测试,同时测试了基于GPU和CPU的计算。

英文摘要

The limited-memory BFGS (L-BFGS) Hessian update scheme is the critical kernel in many quasi-Newton optimization algorithms. The most common approach to implementing L-BFGS uses $2m$ sequential rank-1 updates as part of solving a linear system when there are $m$ history steps. The performance of this approach suffers when the latency of synchronization is significant, and its poor temporal locality increases the memory traffic when vectors do not fit in cache. The compact dense representation of L-BFGS results in an approach that has minimal synchronization latency and better temporal locality, but it requires an additional pass over the basis vectors and an additional basis that must be recomputed when the $B_0$ matrix changes as in variable-metric methods. In the Portable Extensible Toolkit for Scientific Computation and the Toolkit for Advanced Optimization (PETSc/TAO), we have implemented an intermediate dense formulation of BFGS that retains most of the good characteristics of both the recursive and compact dense approaches. We report single-node performance tests of these implementations on the U.S. Department of Energy's Polaris and Frontier machines, testing both GPU-based and CPU-based computations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑