arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过主导子空间波动的小批量噪声降低锐度

Mini-batch Noise Lowers Sharpness via Dominant-Subspace Fluctuations

Junho So, Dongwook Shin

arXiv 2607.23012首次发表:更新:

发表机构

Ajou University(韩国亚洲大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究SGD训练中主导子空间与损失、锐度的关系,通过推导主导方向波动的平均梯度及小批量噪声引起的锐度校正项,实验表明添加校正项能让GD锐度演变更接近SGD。

AI 中文摘要

在随机梯度下降(SGD)训练期间,梯度通常与损失函数海森矩阵的前k个特征向量所跨越的主导子空间强烈对齐。虽然这似乎自然意味着损失减少主要发生在这个空间内,但先前的工作表明,在这个主导子空间内的更新在减少损失方面没有取得有意义的进展。在这项工作中,我们认为主导子空间更好地理解为不是损失减少的主要空间,而是解释小批量SGD锐度动态的关键子空间。为了解释主导子空间在降低前k锐度方面的作用,我们展示了主导方向上波动的平均梯度如何产生锐度校正项,并推导了主导方向上小批量噪声引起的锐度校正项。实验结果表明,将导出的校正项添加到梯度下降(GD)中,使GD的锐度演变更接近SGD的锐度演变。

英文摘要

During SGD training, the gradients often align strongly with the dominant subspace spanned by the top-$k$ eigenvectors of the Hessian of the loss. While this seems to naturally imply that loss reduction mainly occurs within this space, prior work has shown that updates within this dominant subspace make no meaningful progress in reducing the loss. In this work, we argue that the dominant subspace is better understood not as the main space for loss reduction, but as a key subspace for explaining the sharpness dynamics of mini-batch SGD. To explain the role of the dominant subspace in reducing top-$k$ sharpness, we show how the averaged gradient over fluctuations in the dominant directions produces a sharpness correction term, and derive a sharpness correction term induced by mini-batch noise in the dominant directions. Experimental results show that adding the derived correction term to GD brings the sharpness evolution of GD closer to that of SGD.

Comments20 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑