多模态对比学习中的表达能力
Expressivity In Multimodal Contrastive Learning
浏览论文内容
中文总结 AI 辅助
该研究针对多模态对比学习的表达能力展开分析,明确CLIP架构的表达能力随模态数量变化,提出Hadamard-CLIP模型,实现任意数量模态联合分布的通用近似且保留CLIP的检索优势。
中文摘要 AI 辅助
对比学习已成为现代表示学习的基石,支撑着CLIP式模型,这类模型是文本到图像生成、视觉语言模型以及不断扩展的多模态检索任务的基础。尽管取得了这些实证成功,但这类架构的表达能力仍鲜为人知。为深入理解,我们采用总体层面的密度估计视角研究表达能力:每种架构都包含一组参数化密度,其参数可被选择以近似多模态的联合分布。这一过程分离出了纯表示能力的问题:给定的对比参数化族可任意精度近似哪些联合分布?我们表明,表达能力与架构高度相关。对于两种模态,简单的双塔CLIP架构是通用近似器。CLIP的一种自然推广(在三种或更多模态存在时被广泛应用)基于对所有成对相似度求和的损失函数,该损失函数虽无法表示任意联合分布,但我们证明其仍具备足够表达能力以匹配所有成对条件分布。受此差距启发,我们提出Hadamard-CLIP,它在现有编码器顶部添加一个学习到的权重向量,在保留CLIP快速、可预计算嵌入检索能力的同时,恢复了对任意数量模态的联合分布的通用近似能力。
英文摘要
Contrastive learning has become a cornerstone of modern representation learning, powering CLIP-style models that underpin text-to-image generation, vision-language models, and retrieval across a rapidly growing range of modalities. Despite this empirical success, the expressive power of these architectures remains poorly understood. To gain insight, we study expressivity by adopting a population-level, density-estimation viewpoint: each architecture comprises a parameterized set of densities whose parameters may be chosen to approximate the joint distribution of the modalities. This isolates a question of pure representational capacity: which joint distributions can a given contrastive family of parameterizations approximate to arbitrary accuracy? We show that expressivity is sharply architecture-dependent. For two modalities, the simple two-tower CLIP architecture is a universal approximator. A natural generalization of CLIP, widely used in practice when three or more modalities are present, is based on a loss found by summing over all pairwise similarities. This provably cannot represent arbitrary joint distributions, although we prove that it remains expressive enough to match all pairwise conditionals. Motivated by this gap, we propose Hadamard-CLIP, which adds a single learned weight vector on top of the existing encoders and restores universal approximation of the joint for any number of modalities while preserving CLIP's fast, precomputable-embedding retrieval.
发表机构
- California Institute of Technology(加州理工学院)
机构由 AI 辅助整理,请以论文原文为准。