发表机构
Microsoft; Microsoft Research; Brown University(微软; 微软研究院; 布朗大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文发现CLIP中文本编码器过大会导致零样本性能下降,通过模态特定权重衰减可恢复并提升性能,并提出了更高效的配置方案。
AI 中文摘要
对比语言-图像预训练(CLIP)是许多机器学习应用的构建模块。缩放定律指导了大规模训练的资源分配,然而先前的工作将CLIP总模型大小视为单一变量,没有探索编码器之间的容量分配如何影响下游性能。在此,我们训练了多个具有不同视觉和文本编码器大小的CLIP模型,揭示出对于大多数视觉编码器,存在一个最优的文本编码器大小,超过该大小后零样本性能会下降——即使总参数数量增加。利用这一行为,我们可以获得高效的配置,在参数减少多达55%的情况下,匹配标准ViT-B/16架构的零样本性能。我们进一步表明,这种退化源于过大的文本编码器引起的过拟合,而使用模态特定的权重衰减系数不仅能够恢复,还能在所有退化配置上提升性能。几何分析揭示了一种权衡:扩大文本编码器改善了嵌入均匀性,但恶化了跨模态对齐;我们进一步表明这些指标可以预测零样本性能。我们希望这些发现能激励CLIP架构和训练方法去抵消这种退化,这是可靠且高效缩放CLIP的先决条件。
英文摘要
Contrastive Language-Image Pretraining (CLIP) is a building block of many machine learning applications. Scaling laws have guided resource allocation for large-scale training, yet prior work treats total CLIP model size as a single variable, without exploring how the capacity split between encoders impacts downstream performance. Here, we train multiple CLIP models with different vision and text encoder sizes, revealing that for most vision encoders, there is an optimal text encoder size beyond which zero-shot performance degrades---even as total parameter count increases. Exploiting this behavior yields efficient configurations that match the zero-shot performance of the standard ViT-B/16 architecture with up to 55% fewer parameters. We further show that this degradation stems from overfitting induced by the oversized text encoder, and that using modality-specific weight decay coefficients not only recovers but improves performance across all degraded configurations. A geometric analysis reveals a trade-off in which scaling the text encoder improves embedding uniformity but worsens cross-modal alignment; we further show that these metrics are predictive of zero-shot performance. We hope these findings motivate CLIP architectures and training methods that counteract this degradation, a prerequisite for scaling CLIP reliably and efficiently.