arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

关于 Sharpness-Aware Minimization 的隐式平坦性偏差:带定量超参数边界的线性稳定性分析

On the Implicit Flatness Bias of Sharpness-Aware Minimization: A Linear Stability Analysis with Quantitative Hyperparameter Bounds

Jiaxin Deng, Junbiao Pang

arXiv 2608.03197首次发表:更新:

AI 中文总结

该研究针对SAM偏向平坦极小值的隐式偏差,通过线性稳定性分析得到定量超参数边界,验证了增大ρ可降低Hessian最大特征值,并提出TLC-SAM变体进一步优化性能。

AI 中文摘要

Sharpness-Aware Minimization(SAM)通过寻找损失对局部对抗扰动具有鲁棒性的参数来提升泛化性能,但其偏向平坦极小值的隐式偏差背后的定量机制仍不清楚。特别是,扰动半径ρ通常被视为独立的调参,尽管它定义了SAM测量尖锐度的邻域。我们通过线性稳定性分析了插值极小值附近的小批量SAM。在局部线性化和梯度噪声对齐假设下,我们证明每个线性稳定的极小值满足λ_max ≤ 三次根号下(bΓ/(2ρη²)),其中λ_max是Hessian最大特征值,b是批量大小,η是学习率,Γ是梯度范数的边界。该边界定量表征了SAM的隐式平坦性偏差:在其他量固定时,更小的批量大小、更大的学习率或更大的半径会将线性稳定的SAM限制在更平坦的极小值。它还揭示了必要的权衡:ρ应足够大以促进平坦性,但需保持足够局部以保留近似和稳定训练。我们在CIFAR-100上用ResNet-18和VGG-19对900个模型进行受控研究验证了该预测,在所有批量大小和学习率设置下,增大ρ始终与更小的Hessian最大特征值相关。最后,我们将该分析实例化为Taylor-Locality Controlled SAM(TLC-SAM),它利用观测到的Taylor近似误差调整ρ,相比固定半径SAM进一步降低了顶部Hessian特征值。我们的结果为分析和设计SAM变体提供了定量超参数边界以及稳定性-局部性视角。

英文摘要

Sharpness-Aware Minimization (SAM) improves generalization by seeking parameters whose loss is robust to local adversarial perturbations, but the quantitative mechanism underlying its implicit bias toward flat minima remains unclear. In particular, the perturbation radius $ρ$ is typically treated as an isolated tuning parameter, despite defining the neighborhood in which SAM measures sharpness. We analyze mini-batch SAM near an interpolating minimum through linear stability. Under local linearization and gradient-noise alignment assumptions, we prove that every linearly stable minimum satisfies $λ_{\max}\leq\sqrt[3]{bΓ/(2ρη^2)}$, where $λ_{\max}$ is the largest Hessian eigenvalue, $b$ is the batch size, $η$ is the learning rate, and $Γ$ bounds the gradient norm. The bound quantitatively characterizes SAM's implicit flatness bias: holding the other quantities fixed, a smaller batch size, a larger learning rate, or a larger radius restricts linearly stable SAM to flatter minima. It also exposes a necessary trade-off: $ρ$ should be large enough to promote flatness, yet remain local enough to preserve the approximation and stable training. We validate this prediction in a controlled study of 900 models on CIFAR-100 with ResNet-18 and VGG-19, where increasing $ρ$ is consistently associated with a smaller largest Hessian eigenvalue across batch-size and learning-rate settings. Finally, we instantiate the analysis in Taylor-Locality Controlled SAM (TLC-SAM), which adjusts $ρ$ using the observed Taylor-approximation error and further reduces the top Hessian eigenvalue relative to fixed-radius SAM. Our results provide quantitative hyperparameter bounds and a stability--locality perspective for analyzing and designing SAM variants.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑