arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21999cs.LGcs.AI

从扰动校正到几何感知采样:长尾学习中用于平衡平坦最小值的锐度引导平衡采样

From Perturbation Correction to Geometry-Aware Sampling: Sharpness-Guided Equilibrium Sampling for Balanced Flat Minima in Long-Tailed Learning

Jiaxin Deng, Junbiao Pang

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对长尾学习中泛化差的问题,提出锐度引导平衡采样(SGS)方法,通过调整采样分布优化几何,结合累积类别计数和EMA锐度估计动态调整小批次,实验证明能提升准确率且训练时间短,还开辟了损失景观控制新途径。

中文摘要 AI 辅助

长尾学习存在两个导致泛化能力差的问题:头部类别主导训练曝光,而代表性不足的类别往往收敛到损失景观的更尖锐区域。传统重采样解决了前者但未考虑几何,现有的长尾锐度感知最小化(SAM)方法仅在抽取有偏差的小批次后修改损失或扰动。我们引入锐度引导平衡采样(SGS),将采样分布视为优化几何的主动控制变量。SGS仅使用标准SAM更新获得的累积类别计数和EMA锐度估计,通过增加较少采样类别的采样概率并抑制具有大SAM诱导损失变化的类别,动态调整后续小批次,无需逐类扰动或额外反向传播。我们通过连续时间随机微分方程和依赖采样的PAC - Bayes分析表征此采样过程,解释频率 - 锐度反馈如何使训练朝着更平衡的平坦度分布发展。在不平衡率为100的CIFAR - 100 LT上,SGS - SAM在尾部准确率上比Focal - SAM提高10.85分,总体提高3.56分。在ImageNet - LT上,它在尾部类别上比ImbSAM提高6.59分,总体提高1.20分。其训练时间仅为普通SAM的1.02倍。此外,SGS建立了一条损失景观控制的采样侧途径,表明未来长尾方法可联合调节数据曝光和优化几何,而非将两者视为固定。

英文摘要

Long-tailed learning couples two sources of poor generalization: head classes dominate training exposure, while under-represented classes often converge to sharper regions of the loss landscape. Conventional re-sampling addresses the former without considering geometry, whereas existing long-tailed sharpness-aware minimization (SAM) methods modify losses or perturbations only after biased mini-batches have been drawn. We introduce Sharpness-Guided Equilibrium Sampling (SGS), which treats the sampling distribution as an active control variable for optimization geometry. SGS dynamically adjusts subsequent mini-batches by increasing the sampling probability of less frequently sampled classes while suppressing classes with large SAM-induced loss changes, using only cumulative class counts and EMA sharpness estimates obtained from the standard SAM update, without class-wise perturbations or additional backward passes. We characterize this sampling process through a continuous-time stochastic differential equation and a sampling-dependent PAC-Bayes analysis, explaining how frequency-sharpness feedback can move training toward a more balanced flatness profile. On CIFAR-100 LT with an imbalance ratio of 100, SGS-SAM improves Focal-SAM by 10.85 points in tail accuracy and 3.56 points overall. On ImageNet-LT, it improves ImbSAM by 6.59 points on tail classes and 1.20 points overall. Its training time is only $1.02\times$ that of vanilla SAM. Beyond these gains, SGS establishes a sampling-side route to loss-landscape control, suggesting that future long-tailed methods can jointly regulate data exposure and optimization geometry rather than treating either as fixed.

↑