arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

关于多模态表示学习中的模态差距和对比损失

On the modality gap and the contrastive loss in multi-modal representation learning

Fabian Mager, Hiba Nassar, Lars Kai Hansen

arXiv 2607.10698首次发表:更新:

发表机构

Technical University of Denmark(丹麦技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究CLIP风格双编码器对比学习中的模态差距,发现是InfoNCE公式在低温下的模式失败所致。提出xNCE方法,使用跨模态和模态内负对比对,能缩小模态差距,提升零样本分类性能且不牺牲判别几何。

AI 中文摘要

我们研究了CLIP风格的双编码器对比学习中的模态差距,即图像和文本嵌入在共享空间中训练后仍未对齐。我们认为这种差距是由独立编码器的InfoNCE公式失败引起的。我们进行了单模态实验,发现InfoNCE在低温下会主动产生差距。我们对这一现象进行了理论分析,表明模态差距确实是InfoNCE的模式失败,但仅在低温下。我们提出了一种简单的修改xNCE,它使用跨模态和模态内负对比对。xNCE在MS-COCO上匹配检索性能,同时即使在低温下也能持续缩小差距。值得注意的是,xNCE在所有基准上都比InfoNCE基线提高了零样本分类,而高温InfoNCE和正则化InfoNCE都未能做到这一点,表明xNCE在不牺牲转移所需的判别几何的情况下缩小了模态差距。

英文摘要

We study the modality gap in CLIP-style dual-encoder contrastive learning, where image and text embeddings remain misaligned despite being trained in a shared space. We argue that the gap is induced by a failure of the InfoNCE formulation with independent encoders. We conduct a uni-modal experiment with two independent encoders and identical initialization conditions and find that InfoNCE actively generates a gap at low temperatures. We provide a theoretical analysis of this phenomenon and show that the modality gap is indeed a mode-failure of InfoNCE, but only at low temperatures. We propose a simple modification called xNCE, which uses intermodal as well as intra-modality negative contrastive pairs. xNCE matches retrieval performance on MS-COCO while consistently reducing the gap even at low temperatures. Notably, xNCE improves zero-shot classification over the InfoNCE baseline across all benchmarks, whereas high-temperature InfoNCE and regularized InfoNCE both fail to do so, demonstrating that xNCE reduces the modality gap without sacrificing the discriminative geometry needed for transfer.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑