发表机构
Nanjing University of Science and Technology; Alibaba Group(南京理工大学; 阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态优化中置信度不对称问题,提出MaxCR正则化方法,通过动态干预模态语义置信度,提升多模态分类整体性能。
AI 中文摘要
多模态学习(MML)由于模态不平衡现象而陷入优化困境,导致实践中整体性能次优。尽管许多尝试主要侧重于平衡各模态间的优化动态以解决此问题,我们识别出一个微妙但关键的缺陷:优化产生预测确定性的不对称增益,强模态比弱模态更自信,导致模态贡献不平衡。在本文中,我们的分析揭示该缺陷源于单模态特性而非多模态学习,且这种置信度差异可通过正向跨模态干预加以纠正。基于此见解,我们提出多模态最大置信度正则化(MaxCR)以动态干预模态语义置信度。具体而言,利用非线性稀疏度度量跟踪每个模态的语义置信度。然后基于该度量设计最大抑制和最大激发,分别对强模态和弱模态进行正则化。它们分别惩罚和鼓励top-1置信度,从而约束多模态预测。为此,期望强模态和弱模态做出校准的置信度,从而提升整体性能。在广泛使用的数据集上的实证实验通过与各种最先进(SOTA)多模态学习基线的比较揭示了我们方法的优越性。
英文摘要
Multimodal learning (MML) falls into the optimization dilemma due to the modality imbalance phenomenon, leading to suboptimal overall performance in practice. While many attempts primarily focus on balancing the optimization dynamics across modalities to address this issue, we identify a subtle yet critical flaw: optimization yields asymmetric gains in predictive certainty, with the strong modality more confident than the weak one, driving imbalanced modality contributions. In this paper, our analysis reveals that this flaw stems from unimodal characteristics rather than multimodal learning, and this confidence discrepancy can be corrected by positive cross-modal intervention. Based on this insight, we propose multimodal Max Confidence Regularization (MaxCR) to dynamically intervene in modality semantic confidence. Specifically, the semantic confidence of each modality is tracked using a nonlinear sparsity measure. We then design max suppression and max excitation based on this measure to regularize strong and weak modalities, respectively. They penalize and encourage the top-1 confidence, thereby constraining multimodal prediction. To this end, strong and weak modalities are expected to make calibrated confidence, thereby improving the overall performance. Empirical experiments on widely used datasets reveal the superiority of our method through comparison with various state-of-the-art (SOTA) multimodal learning baselines.