解耦熵最小化
Decoupled Entropy Minimization
- HUST AI and Visual Learning Lab (HAIV Lab)(华中科技大学人工智能与视觉学习实验室)
- Huazhong University of Science and Technology (HUST)(华中科技大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文通过将熵最小化解耦为簇聚合驱动因子和梯度缓解校准器,揭示其奖励坍缩与易分类偏差问题,并提出自适应解耦熵最小化AdaDEM,在噪声和动态环境下的不完全监督任务中取得更优性能。
AI中文摘要:
熵最小化(EM)在机器学习中对于减少类别重叠、弥合领域差距以及限制各类任务的不确定性是有益的,但其潜力受到限制。为研究EM的内部机制,我们将经典EM重新表述并解耦为两个具有相反作用的部分:簇聚合驱动因子(CADF)奖励主导类别并促使输出分布呈现尖峰,而梯度缓解校准器(GMC)则基于预测概率惩罚高置信度类别。此外,我们揭示了经典EM因其耦合表述而产生的局限性:1)奖励坍缩阻碍了高确定性样本在学习过程中的贡献,2)易分类偏差导致输出分布与标签分布之间的错位。为解决这些问题,我们提出了自适应解耦熵最小化(AdaDEM),该方法对CADF带来的奖励进行归一化,并采用边际熵校准器(MEC)替代GMC。AdaDEM优于DEM*(经典EM的一个上界变体),并在噪声和动态环境下的各种不完全监督学习任务中取得了优越性能。
英文摘要:
Entropy Minimization (EM) is beneficial to reducing class overlap, bridging domain gap, and restricting uncertainty for various tasks in machine learning, yet its potential is limited. To study the internal mechanism of EM, we reformulate and decouple the classical EM into two parts with opposite effects: cluster aggregation driving factor (CADF) rewards dominant classes and prompts a peaked output distribution, while gradient mitigation calibrator (GMC) penalizes high-confidence classes based on predicted probabilities. Furthermore, we reveal the limitations of classical EM caused by its coupled formulation: 1) reward collapse impedes the contribution of high-certainty samples in the learning process, and 2) easy-class bias induces misalignment between output distribution and label distribution. To address these issues, we propose Adaptive Decoupled Entropy Minimization (AdaDEM), which normalizes the reward brought from CADF and employs a marginal entropy calibrator (MEC) to replace GMC. AdaDEM outperforms DEM*, an upper-bound variant of classical EM, and achieves superior performance across various imperfectly supervised learning tasks in noisy and dynamic environments.