发表机构
School of Computing and Mathematical Sciences, University of Leicester; Department of Physics and Astronomy, University of Leicester; Scientific Computing, Rutherford Appleton Laboratory, Science and Technology Facilities Council; School of Computer Science and Informatics, De Montfort University; School of Automation and Information Engineering, Xi’an University of Technology(莱斯特大学计算与数学科学学院; 莱斯特大学物理与天文学系; 科学技术设施委员会卢瑟福·阿普尔顿实验室科学计算部; 德蒙福特大学计算机科学与信息学院; 西安理工大学自动化与信息工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对不平衡数据集上传统Softmax分类器性能下降问题,提出基于贝叶斯理论的类平衡Softmax(CBS)方法,它是简单的对数几率调整,计算成本低且易集成。CBS可缓解模型偏好问题,实验证明其高度可扩展且优于现有方法。
AI 中文摘要
使用传统Softmax分类器的深度学习模型在各种分类任务中取得了显著成功。然而,它们在不平衡数据集上的性能会显著下降。虽然平衡Softmax被广泛用作一种先进的再平衡方法,但它存在固有局限性,比如对尾部类别的测试准确率异常低。为了缓解这些缺点,我们提出了类平衡Softmax(CBS)。基于理论贝叶斯框架和启发式幂律假设,CBS是一种简单的对数几率调整,计算成本低且易于集成到现有流程中。此外,我们刻画了在不平衡数据上训练的模型中的一个基本现象,即偏好问题,其中模型对数据有限的类表现出更高的训练误差和更大的泛化差距。为了量化这个问题,我们引入了一种新的度量,并证明CBS有效地缓解了偏好问题。在大规模基准上的大量实验表明,CBS具有高度可扩展性,并且优于包括平衡Softmax在内的现有方法。
英文摘要
Models trained on long-tailed data using standard softmax tend to exhibit higher training error and a larger generalisation gap for classes with fewer training samples. We characterise this class-wise disparity as the preference issue and quantify it using a new metric, the model imbalance level $I$. To understand this issue, we analyse how imbalanced training data adversely affects class-wise gradients under standard softmax training. This paper then develops a finite-data Generalised Balanced Softmax (GBS) framework for analysing and mitigating the preference issue. The framework uses the training-time logit adjustment $z_{nc}+β\log|N_c|$, which is algebraically identical to the training-time logit-adjusted loss of Menon et al. (2021) when $τ=β$. The case $β=1$ also coincides with Balanced Softmax and with the unit adjustment supported by the Fisher-consistency argument under the true data distribution, corresponding to an idealised infinite-data setting. Building on this existing loss family, this paper uses a heuristic power-law assumption to motivate the adjustable coefficient and studies how $β$ affects trained models. Across the evaluated long-tailed benchmarks, $β=1$ does not attain the highest average testing recall on most datasets, showing that a different coefficient can be preferable when training on finite data. The selected values of $β$ reduce $I$ and improve average testing recall relative to the $β=1$ reference, while retaining negligible computational overhead and compatibility with existing representation-learning frameworks.
Comments26 pages, 8 figures. Revised author version of the CBS work published in Pattern Recognition. Title and terminology updated to Generalised Balanced Softmax (GBS); theoretical presentation and relationship to existing logit-adjustment methods clarified. Underlying loss function and reported experimental results unchanged
Journal refPattern Recognition (2026), Article 114961
DOI:10.1016/j.patcog.2026.114961