基于局部曲率和锐度感知最小化的f-散度正则化解释
Explaining f-Divergence-Based Regularization via Local Curvature and Sharpness-Aware Minimization
浏览论文内容
中文总结 AI 辅助
本文通过局部曲率分析揭示f-散度正则化与锐度感知最小化在参数和输入扰动下的等价性,并实验证明最大曲率惩罚可提升泛化性能。
中文摘要 AI 辅助
基于散度的正则化和锐度感知最小化(SAM)是深度学习中改善泛化能力的两种重要方法,二者均源于对扰动的鲁棒性动机。然而,它们之间的关系在很大程度上尚未被探索。基于f-散度的经典二阶展开,我们证明这两种方法在参数空间扰动下是局部一致的:两者都引入曲率敏感惩罚,其中散度正则化产生Fisher加权二次型,而SAM通过主导Hessian特征值惩罚锐度。对于具有指数族输出分布的负对数似然目标,这种对应关系变得尤为透明,因为Fisher矩阵和Gauss-Newton矩阵一致。我们进一步表明,相同的局部几何视角可扩展到输入空间扰动,其中基于散度的正则化通过输入的变换来定义。在此设置下,正则化器在输入空间上诱导拉回二次型,提供了比标准SAM更通用的扰动框架,同时保持相同的局部敏感性解释。为实证验证该分析,我们使用非对称α-偏斜Jensen-Shannon散度(JSD)族作为受控测试平台。其局部曲率系数按α(1-α)缩放,并在对称点α=1/2处最大化,此时恢复标准JSD。输入扰动机制下的损失景观可视化表明,更强的诱导曲率惩罚与更平坦的局部最小值相关。在四个基准数据集上的实验进一步表明,准确率和负对数似然在最大曲率惩罚附近区域始终表现最佳。
英文摘要
Divergence-based regularization and Sharpness-Aware Minimization (SAM) are two prominent approaches for improving generalization in deep learning, both motivated by robustness to perturbations. However, their relationship has remained largely unexplored. Building on classical second-order expansions of $f$-divergences, we show that the two methods are locally consistent under parameter-space perturbations: both induce curvature-sensitive penalties, with divergence regularization yielding a Fisher-weighted quadratic form and SAM penalizing sharpness through the dominant Hessian eigenvalue. For negative log-likelihood objectives with exponential-family output distributions, this correspondence becomes especially transparent, since the Fisher and Gauss-Newton matrices coincide. We further show that the same local geometric perspective extends to input-space perturbations, where divergence-based regularization is defined through transformations of the input. In this setting, the regularizer induces a pullback quadratic form on the input space, providing a more general perturbation framework than standard SAM while preserving the same local sensitivity interpretation. To validate the analysis empirically, we use the asymmetric $α$-skew Jensen-Shannon divergence (JSD) family as a controlled testbed. Its local curvature coefficient scales as $α(1-α)$ and is maximized at the symmetric point $α=\tfrac12$, which recovers the standard JSD. Loss-landscape visualizations in the input-perturbation regime show that stronger induced curvature penalization is associated with flatter local minima. Experiments on four benchmark datasets further demonstrate that both accuracy and negative log-likelihood are consistently best near this regime of maximal curvature penalization.
发表机构
- EURECOM
- University of Granada(格拉纳达大学)
机构由 AI 辅助整理,请以论文原文为准。