arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新思考CLIP后训练中的对比损失:带有冻结文本编码器的互补框架

Rethinking Contrastive Loss in CLIP Post-training: A Complementary Framework with Frozen Text Encoder

Zidan Wang, Yaqian Li, Xiaokai Zhang, Kaiwen Long, Kun He, Hanpeng Liu

arXiv 2610.11374首次发表:更新:

发表机构

Huazhong University of Science and Technology; Li Auto Inc.(华中科技大学; 理想汽车)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出ComCLIP框架,通过冻结CLIP文本编码器、采用合适温度的对比损失等优化,提升了CLIP视觉编码器的特征可迁移性,且作为即插即用组件适配LLaVA等下游VLM。

AI 中文摘要

CLIP是基础的视觉-语言模型,也是LLaVA等下游视觉-语言模型(VLM)事实上的视觉编码器。后训练为优化CLIP提供了轻量途径,但近期研究指出,标准对比损失因小批量下的灾难性遗忘不适用于后训练,促使研究者放弃对比目标转而采用蒸馏方案。本文重新审视这一前提,发现InfoNCE目标下,已报道的遗忘主要并非由负样本不足导致,而是由对比温度τ的取值不当所致:当τ设置得足够小时,对比后训练会提升而非降低预训练CLIP的性能,本文通过InfoNCE梯度的温度依赖性对此作出解释。基于该发现,本文提出ComCLIP,这是一种轻量的单轮后训练方案,它冻结CLIP的文本编码器,因此优化后的视觉编码器可作为即插即用替换项,架构与推理成本均保持不变;该方案通过设置合适温度的对比损失、针对原始CLIP的MSE锚定损失,以及来自DINOv2的关系蒸馏损失来训练视觉编码器。在多个随机种子下,ComCLIP在零样本分类上与自蒸馏基线CLIP-Refine表现相当,同时显著提升视觉特征的可迁移性——以线性探测衡量,ViT-B/16上的结果为48.99,而CLIP-Refine为42.28;在ViT-L/14上,其MMVP指标较CLIP-Refine提升(24.20 vs. 19.01),不过CLIP-Refine在图像-文本检索上仍更具优势。将ComCLIP作为即插即用视觉编码器用于LLaVA-1.5-7B,无需重新调整投影器或大语言模型(LLM),在8项VLM基准测试中未产生净变化,即该优化不会破坏下游兼容性。代码与模型可通过指定URL获取。

英文摘要

CLIP serves as a foundational vision-language model and the de facto vision encoder for downstream VLMs such as LLaVA. Post-training offers a lightweight route to refine CLIP, but recent work argues that the standard contrastive loss is unsuitable for post-training due to catastrophic forgetting under small batches, motivating designs that abandon the contrastive objective in favor of distillation. We revisit this premise and find that, for the InfoNCE objective, the reported forgetting is driven primarily not by insufficient negatives but by an inappropriate magnitude of the contrastive temperature $τ$: with $τ$ set sufficiently small, contrastive post-training improves rather than degrades the pretrained CLIP, which we explain through the temperature dependence of the InfoNCE gradient. Building on this finding, we propose \textbf{ComCLIP}, a lightweight single-epoch post-training recipe that freezes CLIP's text encoder---so the refined vision encoder is a drop-in replacement with unchanged architecture and inference cost---and trains the vision encoder with a properly-tempered contrastive loss, an MSE anchoring loss against the original CLIP, and a relational distillation loss from DINOv2. Over multiple seeds, ComCLIP matches the self-distillation baseline CLIP-Refine on zero-shot classification while significantly improving the transferability of visual features, measured by linear probing ($48.99$ vs.\ $42.28$ on ViT-B/16), and on ViT-L/14 it also improves MMVP over CLIP-Refine ($24.20$ vs.\ $19.01$); CLIP-Refine remains stronger on image-text retrieval. Used as a drop-in vision encoder for LLaVA-1.5-7B without re-aligning the projector or LLM, ComCLIP yields no net change across $8$ VLM benchmarks, i.e., the refinement does not break downstream compatibility. Code and models are available at https://github.com/showstarpro/ComCLIP.git.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑