发表机构
Berlin Institute of Health, Charité - Universitätsmedizin Berlin; Department of Mathematics and Computer Science, Freie Universität Berlin; Intelligent Medicine Institute, Fudan University(柏林健康研究院,柏林夏里特大学医学中心; 柏林自由大学数学与计算机科学系; 复旦大学智能医学研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究多模态对比学习从成对扩展到高阶对齐的挑战,确定编码器雅可比条件是关键因素,引入几何保留编码器,通过正则化调节雅可比,实验证明改善其条件可提升检索和线性探针性能,表明多模态对比学习还取决于编码器几何和优化属性。
AI 中文摘要
对比学习正越来越多地转向具有三种或更多模态的设置,而非图像-文本对。然而,将模型从成对扩展到高阶多模态对齐会带来优化和表示方面的挑战。我们确定编码器雅可比条件是三模态对比学习的关键因素:条件不佳的编码器会出现奇异值谱坍塌或放大,导致雅可比条件数爆炸和多模态对齐退化。我们通过正则化直接调节雅可比来引入几何保留编码器(GPE),并证明像LeakyReLU激活和残差路径这样的简单修改能恢复这些几何益处。在包括缺失模态的合成基准和四个真实世界数据集上,改善雅可比条件能提升多个对比目标的检索和线性探针性能,而在线性探针中,表达性目标益处不大。更广泛地说,我们的结果表明多模态对比学习不仅取决于目标表现力,还取决于底层编码器的几何和优化属性。
英文摘要
Contrastive learning is increasingly moving toward settings with three or more modalities instead of image-text pairs. Yet, extending models from pairwise to higher-order multimodal alignment can introduce optimization and representation challenges. We identify encoder Jacobian conditioning as a key factor in trimodal contrastive learning: poorly conditioned encoders exhibit collapsing or amplified singular-value spectra, leading to exploding Jacobian condition numbers and degraded multimodal alignment. We introduce geometry-preserving encoders (GPEs) by directly conditioning the Jacobian through regularization and demonstrating that simple modifications like LeakyReLU activations and residual paths recover these geometric benefits. Across a synthetic benchmark and four real-world datasets including missing modalities, improving Jacobian conditioning boosts retrieval and linear probe performance across multiple contrastive objectives, whereas expressive objectives yield little benefit in linear probes. More broadly, our results show that multimodal contrastive learning depends not only on objective expressivity, but also on the geometric and optimization properties of the underlying encoders.