arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.19544stat.MLcs.LGmath.OC

RELTA-SGLD:非凸随机梯度朗之万学习的相对增长局部驯服

RELTA-SGLD: Relative-Growth Localized Taming for Nonconvex Stochastic-Gradient Langevin Learning

  • School of Mathematics and Statistics, Yunnan University(云南大学数学与统计学学院)

机构由 AI 辅助整理,请以论文原文为准。

Yiwei Zhou, Ziheng Chen

AI总结:

研究针对非凸随机梯度朗之万学习,提出RELTA-SGLD驯服方案,通过阈值和相对增长原理稳定更新,减少不必要抑制,证明了相关稳定性和精度,在实验中表现良好,改进了学习指标并保持学习动态。

AI中文摘要:

我们引入了RELTA-SGLD,这是一种驯服方案,可稳定超线性随机梯度更新,同时减少对原始学习漂移的不必要抑制。一个阈值决定驯服开启的位置,而从一步李雅普诺夫稳定性条件导出的相对增长原理决定所需的驯服强度。它们共同产生更轻的λ尺度分母并保持非零远尾回报。结果,我们证明了具有超线性增长随机梯度预言机的非凸SGLD在W1和W2中的多项式矩稳定性和一阶平稳精度,改进了可比随机梯度驯服方案的相应半阶和四分之一阶界。在主动稳定压力下的Fashion-MNIST上,RELTA在平均学习指标上优于未驯服的SGLD和TUSLA,并且与调优后的AdamW参考具有竞争力。在普通训练模式下,其更轻的局部分母减少了对原始更新的不必要扰动,并保持了几乎未驯服的学习动态。

英文摘要:

We introduce RELTA-SGLD, a taming scheme that stabilizes superlinear stochastic-gradient updates while reducing unnecessary suppression of the original learning drift. A threshold determines where the taming turns on, while a relative-growth principle derived from the one-step Lyapunov stability condition determines the required taming strength. Together, they produce a lighter $λ$-scale denominator and preserve a nonvanishing far-tail return. As a consequence, we prove polynomial moment stability and first-order stationary accuracy in both $W_1$ and $W_2$ for nonconvex SGLD with superlinearly growing stochastic-gradient oracles, improving the corresponding half-order and quarter-order bounds for comparable stochastic-gradient tamed schemes. On Fashion-MNIST under active stabilization pressure, RELTA improves the mean learning metrics over both untamed SGLD and TUSLA and remains competitive with a tuned AdamW reference. In an ordinary-training regime, its lighter localized denominator reduces unnecessary perturbation of the original update and maintains nearly untamed learning dynamics.

补充信息

↑