TeMo:多模态对比学习的温度调制
TeMo: Temperature Modulation for Multimodal Contrastive Learning
浏览论文内容
中文总结 AI 辅助
本文提出TeMo温度调制框架,根据样本对相似度自适应调整温度,集成多模态与单模态损失,在零样本检索和分类任务上取得新最优性能。
中文摘要 AI 辅助
对比学习方法通过训练模型使相似样本靠近、不相似样本远离,从而取得优异性能。对比学习的一个关键组成部分是温度超参数 $\ au$,它控制对负样本施加的惩罚强度。然而,大多数现有方法要么固定该超参数,要么在训练过程中学习一个全局值。在本文中,我们提出了TeMo(温度调制框架),一种基于相似度的调制方法,该方法根据每个正负样本对的相似度自适应地调整温度,从而实现更细粒度的多模态对比学习。我们的方法通过逐步过渡,将温度调制的多模态和单模态损失与标准多模态对比损失无缝集成。这种设计使模型能够在不同训练阶段捕获粗粒度和细粒度的语义。大量实验表明,TeMo的每个组件在多种零样本检索和分类任务中均能持续提升性能,并取得了新的最先进结果。
英文摘要
Contrastive learning approaches achieve strong performance by training models to bring similar samples closer while pushing dissimilar samples apart. A crucial component of contrastive learning is the temperature hyperparameter $τ$, which controls the penalty strength applied to negative samples. However, most existing methods either fix this hyperparameter or learn a global value during training. In this paper, we introduce TeMo, Temperature Modulation framework, a similarity-based modulation approach that adaptively adjusts the temperature for each positive-negative pair according to their similarity, enabling more fine-grained multimodal contrastive learning. Our approach seamlessly integrates temperature-modulated multimodal and unimodal losses with the standard multimodal contrastive loss by gradually transitioning between them. This design allows the model to capture both coarse- and fine-grained semantics at different training stages. Extensive experiments demonstrate that each component of TeMo consistently enhances performance across diverse zero-shot retrieval and classification tasks, establishing new state-of-the-art results.
发表机构
- MPI for Informatics(马克斯·普朗克信息学研究所)
- SIC Tuebingen AI Center(图宾根人工智能中心SIC)
- University of Tuebingen(图宾根大学)
- MIT-IBM Watson AI Lab(麻省理工学院-IBM沃森人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。