EMASAM:基于EMA引导扰动的计算高效型锐度感知最小化方法
EMASAM: a Computationally Efficient Sharpness-Aware Minimization via EMA-Guided Perturbations
- Institute of Science Tokyo(东京科学大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出EMASAM,一种无需额外梯度计算的SAM高效变体,通过主模型与EMA影子模型的差异生成扰动,缓解梯度不稳定性,实验验证其效率与鲁棒性。
AI中文摘要:
近期优化研究进展表明,损失景观的锐度是缩小泛化差距的关键因素。受此启发,锐度感知最小化(SAM)作为一种提升泛化能力的训练策略被提出。尽管SAM性能优异,但由于其核心算法在扰动步骤需额外计算梯度,导致计算成本翻倍。为克服这一局限,本文提出指数移动平均锐度感知最小化(EMASAM),它是SAM的计算高效变体。EMASAM在扰动步骤无需损失梯度,而是基于主模型与EMA影子模型的差异定义扰动方向,该扰动从稳定的平均位置向欠稳定区域移动,是SAM最坏情况扰动的更柔和、更廉价替代方案。此外,由于EMASAM的扰动不依赖含噪小批量梯度,它缓解了SAM固有的梯度诱导不稳定性。因此,EMASAM无需额外反向传播,同时保留了SAM式训练的泛化能力。多项实验证实了该方法的效率与鲁棒性。
英文摘要:
Recent progress in optimization research has highlighted the sharpness of the loss landscape as a key factor in narrowing the generalization gap. Motivated by this insight, Sharpness-Aware Minimization (SAM) was proposed as a training strategy that enhances generalization. Despite the promising performance, SAM suffers from its twice computational cost due to its core algorithm requiring an extra gradient computation during the perturbation step. To overcome this limitation, we introduce Exponential Moving Average Sharpness-Aware Minimization (EMASAM), a computationally efficient variant of SAM. EMASAM does not require the loss gradient in the perturbation step. Instead, EMASAM defines the perturbation direction based on the discrepancy between the main model and the EMA shadow model. This perturbation travels away from the stable average position toward the less stable area, acting as a softer yet cheaper alternative to SAM's worst-case scenario perturbation. Moreover, since EMASAM's perturbation does not rely on noisy mini-batch gradients, it mitigates the gradient-induced instability inherent in SAM. Hence, EMASAM eliminates the need for an extra backpropagation while also preserving the generalization ability of the SAM-style training. Several experiments have been performed and confirm the efficiency and robustness of our method.