arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

噪声监督下多模态学习中信任对象的学习

Learning What to Trust in Multimodal Learning under Noisy Supervision

Jiashuo Zou, Xiaobo Xia

arXiv 2610.11057首次发表:更新:

发表机构

University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对实际场景中多模态学习的噪声标签问题,提出多模态标签噪声检测框架REFINE,通过分析表示结构与噪声检测能力的关系,联合融合与单模态表示检测噪声,提升模型泛化性能。

AI 中文摘要

多模态分类通过处理和关联来自多种模态的信息以实现更准确的预测。然而,现有方法通常依赖高质量的真实标签,而在实际场景中这类标签难以获取。尽管用于带噪声标签学习的样本选择方法旨在从带噪声数据中识别标注正确的样本,但传统方法主要关注单模态设置,未能充分利用多模态信息。这促使我们在多模态学习中构建更可靠的噪声检测器。为此,我们从理论上分析了表示结构与噪声检测能力之间的关系。基于该分析,我们提出了REFINE,这是一种多模态标签噪声检测框架,它联合使用融合表示与单模态表示进行标签噪声检测。具体而言,REFINE通过对目标类和背景类进行判别分析来构建判别特征向量,并为每个类别选择具有更强噪声检测能力的可信表示空间。在每个可信空间内,REFINE测量每个样本表示与判别特征向量之间的对齐程度,随后合并从这些空间中选出的子集。合并后的子集提供了更纯净的监督信号,用于更新多模态分类器,从而减少训练过程中标注错误样本的影响,提升模型的泛化能力。在不同任务上开展的大量实验表明,REFINE相较于基线方法具有优越性,其源代码将公开提供。

英文摘要

Multimodal classification processes and relates information from multiple modalities to achieve more accurate predictions. However, existing methods typically rely on high-quality ground-truth labels, which are difficult to obtain in real-world scenarios. While sample-selection methods for learning with noisy labels aim to identify correctly labeled examples from noisy data, traditional methods primarily focus on unimodal settings and fail to exploit multimodal information fully. This motivates us to build a more reliable noise detector in multimodal learning. To this end, we theoretically analyze the relationship between representation structure and noise detection capability. Based on this analysis, we propose REFINE, which is a multimodal label-noise detection framework that jointly uses fused and unimodal representations for label-noise detection. Specifically, REFINE constructs discriminative eigenvectors through discriminative analysis of the target and background classes and selects trusted representation spaces with better noise detection capability for each class. Within each trusted space, REFINE measures the alignment between each instance representation and the discriminative eigenvectors. It then combines the subsets selected from these spaces. The combined set provides cleaner supervision for updating the multimodal classifier, thereby reducing the influence of mislabeled examples during training and improving model generalization. Extensive experiments across diverse tasks demonstrate REFINE's superiority compared to baseline methods. The source code will be publicly available.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑