arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12009math.OCcs.LG

带稳定Barzilai–Borwein步长的自适应Bregman近端随机梯度方法

Adaptive Bregman Proximal Stochastic Gradient with a Stabilized Barzilai--Borwein Step Size

Chenhan Jin, Shengze Xu, Binghui Xie, Kaiwen Zhou, Fan Jia, James Cheng, Tieyong Zeng

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出无需线搜索的Ada-BPSG方法,结合SAGA梯度表与稳定BB步长,在凸、非凸等场景下获收敛速率保证,实验中对初始步长敏感性低且性能优于基线。

中文摘要 AI 辅助

Bregman近端随机梯度(BPSG)方法可将降方差复合优化应用于欧氏平滑性难以良好刻画其几何特性的目标函数。然而,该方法的性能仍对步长敏感:原始随机曲率估计值可能大幅波动,而线搜索会增加多次近端评估。我们提出Ada-BPSG,这是一种无需线搜索的BPSG方法,它将SAGA梯度表与稳定Barzilai–Borwein(BB)候选值相结合。中位数聚合增量正割信息,使几乎奇异的局部比率权重极低;显式保护机制将所得曲率估计转换为收敛所需的有界步长序列。该设计在相对平滑性和分量级方差控制与有限维赋范空间的收敛性之间建立了直接分析链。我们证明,对于凸目标函数,其遍历速率为O(n/K);在相对二次增长条件下,重启后可达到线性速率;对于非凸情形,Bregman近端残差的界为O(1/K)。在逻辑回归和稀疏非负矩阵分解任务上,Ada-BPSG不仅能达到较低的目标函数值,且对初始步长的敏感性远低于标准降方差基线方法,同时避免了线搜索。

英文摘要

Bregman proximal stochastic gradient (BPSG) methods bring variance-reduced composite optimization to objectives whose geometry is poorly captured by Euclidean smoothness. Their performance, however, remains sensitive to the step size: raw stochastic curvature estimates can fluctuate sharply, whereas line searches add repeated proximal evaluations. We introduce Ada-BPSG, a line-search-free BPSG method that couples the SAGA gradient table with a stabilized Barzilai--Borwein (BB) candidate. A mediant aggregates incremental secant information so that nearly singular local ratios receive little weight, and an explicit safeguard translates the resulting curvature estimate into the bounded step-size sequence required for convergence. This design yields a direct analytical chain from relative smoothness and component-wise variance control to convergence in finite-dimensional normed spaces. We prove an $O(n/K)$ ergodic rate for convex objectives, a restarted linear rate under relative quadratic growth, and an $O(1/K)$ bound for a Bregman proximal residual in the nonconvex setting. On logistic regression and sparse nonnegative matrix factorization, Ada-BPSG combines low objective values with substantially less sensitivity to the initial step size than standard variance-reduced baselines, while avoiding line search.

发表机构

  • University of Utah(犹他大学)
  • Beijing Normal-Hong Kong Baptist University(北京师范大学-香港浸会大学联合国际学院)
  • Guangzhou Nanfang College(广州南方学院)
  • The Chinese University of Hong Kong(香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

↑