发表机构
Technical University of Denmark(丹麦技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究CLIP风格双编码器对比学习中的模态差距,发现是InfoNCE公式在低温下的模式失败所致。提出xNCE方法,使用跨模态和模态内负对比对,能缩小模态差距,提升零样本分类性能且不牺牲判别几何。
AI 中文摘要
我们研究了CLIP风格的双编码器对比学习中的模态差距,即图像和文本嵌入在共享空间中训练后仍未对齐。我们认为这种差距是由独立编码器的InfoNCE公式失败引起的。我们进行了单模态实验,发现InfoNCE在低温下会主动产生差距。我们对这一现象进行了理论分析,表明模态差距确实是InfoNCE的模式失败,但仅在低温下。我们提出了一种简单的修改xNCE,它使用跨模态和模态内负对比对。xNCE在MS-COCO上匹配检索性能,同时即使在低温下也能持续缩小差距。值得注意的是,xNCE在所有基准上都比InfoNCE基线提高了零样本分类,而高温InfoNCE和正则化InfoNCE都未能做到这一点,表明xNCE在不牺牲转移所需的判别几何的情况下缩小了模态差距。
英文摘要
We study the modality gap in CLIP-style dual-encoder contrastive learning, where image and text embeddings remain misaligned despite being trained in a shared space. We argue that the gap is induced by a failure of the InfoNCE formulation with independent encoders. We conduct a uni-modal experiment with two independent encoders and identical initialization conditions and find that InfoNCE actively generates a gap at low temperatures. We provide a theoretical analysis of this phenomenon and show that the modality gap is indeed a mode-failure of InfoNCE, but only at low temperatures. We propose a simple modification called xNCE, which uses intermodal as well as intra-modality negative contrastive pairs. xNCE matches retrieval performance on MS-COCO while consistently reducing the gap even at low temperatures. Notably, xNCE improves zero-shot classification over the InfoNCE baseline across all benchmarks, whereas high-temperature InfoNCE and regularized InfoNCE both fail to do so, demonstrating that xNCE reduces the modality gap without sacrificing the discriminative geometry needed for transfer.