arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13569stat.MLcs.LG

一种带动量与自适应迁移率的拉回校正标量辅助变量优化器

A pullback-corrected scalar auxiliary variable optimizer with momentum and adaptive mobility

  • Purdue University(普渡大学)
  • University of Washington(华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

Jiahao Zhang, Shiheng Zhang, Guang Lin

AI总结:

本文提出PB-SAV优化器,通过拉回校正和自适应迁移率处理多分量目标,在Burgers测试中四分量较单分量显著降低目标与误差。

AI中文摘要:

科学机器学习中的目标函数通常被表示为若干项之和,例如物理信息神经网络的残差、边界、初始和数据损失。在拉回校正标量辅助变量(PB--SAV)方法中,一个标量跟踪平移后的目标函数,而各分量的梯度构建一个秩至多为分量数的半正定曲率校正。我们将该校正引入到带动量和自适应迁移率的优化器中,在单个隐式求解中将其应用于梯度和存储的动量。在Loewner序下非增的迁移率可导出精确的修正能量律,涵盖欧几里得型和AMSGrad型选择;求解后附加的动量对应的恒等式包含一个不定符号的交叉项。对于固定迁移率,我们给出了在驻点处局部稳定的充要条件,该条件依赖于Hessian减去两倍校正,并证明该条件对每个标量松弛序列也给出局部几何收敛。隐式求解简化为一个阶数为分量数的稠密系统。在前向Burgers比较中,四个分量相对于一个分量在相同学习率和动量设置下,将平均尾部目标降低了64.7%,最终解误差降低了50.2%。

英文摘要:

Objectives in scientific machine learning are often prescribed as a sum of several terms, such as the residual, boundary, initial, and data losses of a physics-informed neural network. In the pullback-corrected scalar auxiliary variable (PB--SAV) method, one scalar tracks the shifted objective while the component gradients build a positive semidefinite curvature correction of rank at most the number of components. We carry that correction into an optimizer with momentum and an adaptive mobility, applying it to the gradient and the stored momentum in a single implicit solve. A mobility that is nonincreasing in the Loewner order yields an exact modified energy law, covering Euclidean and AMSGrad-type choices; the corresponding identity for momentum appended after the solve carries a cross term of indefinite sign. For a fixed mobility we give a necessary and sufficient condition for local stability at a stationary point, depending on the Hessian minus twice the correction, and show that it also gives local geometric convergence for every scalar relaxation sequence. The implicit solve reduces to a dense system whose order is the number of components. In the forward Burgers comparison, four components reduce the mean tail objective by 64.7% and the final solution error by 50.2% relative to one component at the same learning rate and momentum settings.

↑