arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

具有最优失败指数的随机次梯度方法

A stochastic subgradient method with optimal failure exponent

Bart P. G. van Parys

arXiv 2609.37425首次发表:更新:

发表机构

CWI Amsterdam(荷兰数学与计算机科学研究中心阿姆斯特丹分部)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对随机次梯度优化,提出调和均匀平均调度并证明其达到最优失败指数,且由匹配的不可行性结果保证最优性,小噪声极限退化为对抗性误差模型。

AI 中文摘要

固定目标精度 $\varepsilon$、梯度噪声水平 $s$ 和视界 $N$。我们希望设计能够最小化观察到超过目标精度的次优性间隙的概率的算法,即 $\mathcal{E}_N = -\log \sup_{f,P} \mathbb{P}_P(f(x_A) - f_\star \ge \varepsilon)$,其中噪声律 $P$ 仅已知为次高斯。我们特别提出一种均匀平均的调和调度(形式为 $h_k = R^2/(\varepsilon (N+m-k))$),并通过优化的指数超鞅论证证明其达到最优指数 $\mathcal{E}_N^\star = \varepsilon^2 N (1+o(1))/(2R^2 s^2)$。最优性由一个匹配的不可行性结果保证:在高斯噪声下,梯度掩蔽的测度变换将每个算法的指数限制在同一主阶。在小噪声极限下,我们的设置退化为 Gösgens 和 van Parys (2025) 针对次梯度方法的对抗性误差模型。

英文摘要

Fix a target accuracy $\varepsilon$, a gradient-noise level $s$, and a horizon $N$. We wish to design algorithms which minimize the probability of observing a suboptimality gap which exceeds the target accuracy, i.e., $\mathcal{E}_N = -\log \sup_{f,P} \mathbb{P}_P(f(x_A) - f_\star \ge \varepsilon)$, with the noise law $P$ known only to be sub-Gaussian. We single out a uniformly averaged schedule which is harmonic (of the form $h_k = R^2/(\varepsilon (N+m-k))$) and prove, via an optimized exponential supermartingale argument, that it attains the optimal exponent $\mathcal{E}_N^\star = \varepsilon^2 N (1+o(1))/(2R^2 s^2)$. Optimality is certified by a matching impossibility result: under Gaussian noise, a gradient-masking change of measure caps the exponent of every algorithm at the same leading order. In a small-noise limit, our setting degenerates into the adversarial-error model of Gösgens and van Parys (2025) for subgradient methods.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑