AI 中文总结
研究持续状态依赖偏差下随机梯度下降收敛性,提出残差学习算法,将原任务转为双层优化问题,证明该算法能找到无偏差优化问题的解,量化响应函数对收敛复杂度的影响,数值模拟验证算法有效性。
AI 中文摘要
本文研究了在实现更新存在持续且状态依赖偏差时随机梯度下降(SGD)的收敛性,其中期望更新按响应函数逐分量缩放。首先证明此设置下SGD隐含地优化了一个惩罚问题,其最小值与真实最小值不一致。为缓解收敛失败,将原任务重新表述为等效的双层优化问题并提出基于梯度的残差学习算法。理论分析表明残差学习能找到原无偏差优化问题的解。还通过硬件条件数量化响应函数对收敛复杂度的影响,通过构造硬实例表明一般对其多项式依赖不可避免。数值模拟支持了理论结果。
英文摘要
This paper studies the convergence of stochastic gradient descent (SGD) when the implemented updates are subject to a persistent and state-dependent bias, in which the desired update is scaled by response functions component-wise. Our first contribution is to demonstrate that SGD in this setting implicitly optimizes a penalized problem whose minimizer does not coincide with the true minimizer. To mitigate this convergence failure, we reformulate the original task as an equivalent bilevel optimization problem and propose a gradient-based algorithm, termed Residual Learning. Theoretical analysis shows that Residual Learning finds a solution to the original, unbiased optimization problem despite the hardware imperfections. Beyond exact convergence, we quantify how the response functions affect convergence complexity via the hardware condition number and show that a polynomial dependence on it is unavoidable in general, via a construction of a hard instance. The theoretical results are supported by numerical simulations that demonstrate the effectiveness of the proposed algorithm.