arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

锐度感知最小化与Muon:谱范数下的鲁棒性

Sharpness-Aware Minimization and Muon: Robustness under the Spectral Norm

Wenzhi Zhong, Edward Milsom, Michael Murray

arXiv 2607.26001首次发表:更新:

发表机构

University of Bath(巴斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究如何通过锐度感知最小化(SAM)提高泛化能力,针对其“小”扰动依赖几何结构的问题,在SAM两阶段引入矩阵感知几何结构,结合逐层谱内扰动与Muon等外更新,实验表明该组合在ViT-Small/16和ResNet-50上验证准确率最佳。

AI 中文摘要

锐度感知最小化(SAM)旨在通过鼓励对小的最坏情况参数扰动不敏感来提高泛化能力。然而,“小”扰动的概念本质上依赖于几何结构:虽然现有的SAM变体探索了广泛的选择,但对于在实践中哪种几何结构最有效仍缺乏清晰的认识。最近关于矩阵感知优化的工作,特别是Muon优化器,表明尊重隐藏层权重的矩阵结构可以带来强大的经验性能。受此启发,我们在SAM的两个阶段研究矩阵感知几何结构:我们为矩阵值隐藏层参数引入了逐层谱内扰动,并将其与AdamW/SGDW或Muon在外部更新中相结合。在ImageNet-1K上对ViT-Small/16和ResNet-50进行的实验中,我们发现谱内步长与Muon外步长的组合始终表现强劲,在所评估的方法中在两个模型上都实现了最佳的验证准确率。

英文摘要

Sharpness-Aware Minimization (SAM) aims to improve generalization by encouraging insensitivity to small, worst-case parameter perturbations. However, the notion of a "small" perturbation is inherently geometry-dependent: while existing SAM variants have explored a wide range of choices, a clear perspective on which geometries are most effective in practice remains elusive. Recent work on matrix-aware optimization, particularly the Muon optimizer, suggests that respecting the matrix structure of hidden-layer weights can lead to strong empirical performance. Motivated by this, we study matrix-aware geometry in both stages of SAM: we introduce a layerwise spectral inner perturbation for matrix-valued hidden-layer parameters and combine it with either AdamW/SGDW or Muon in the outer update. Across ImageNet-1K experiments on ViT-Small/16 and ResNet-50, we find that the combination of a spectral inner step with a Muon outer step performs consistently strongly, achieving the best validation accuracy on both models among the evaluated methods.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑