发表机构
Tsinghua University; Cisco Research; University of Illinois Chicago(清华大学; 思科研究院; 伊利诺伊大学芝加哥分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究 CLIP 潜在空间,提出基于冯·米塞斯-费舍尔分布混合的密度模型,用期望最大化算法学习,实现准确可解释的密度估计,改善长尾和分布外检测,建立几何一致的概率框架。
AI 中文摘要
对比语言-图像预训练(CLIP)表示形成了一个由余弦相似度控制的语义嵌入空间,反映了内在的超球面几何。然而,现有的概率解释通常依赖高斯假设,无法捕捉这种方向性和多模态结构。我们基于在单位超球面上定义的冯·米塞斯-费舍尔(MovMF)分布混合,为 CLIP 潜在空间提出了一个有原则的密度模型。使用期望最大化(EM)算法,我们有效地学习了一个概率模型,其中每个混合成分对应一个连贯的语义概念。该公式产生了一个与超球面几何自然对齐的闭式似然,实现了准确且可解释的密度估计。实验上,我们的模型显著改善了长尾和分布外检测,并提供了自然的语义分解,将每个嵌入表示为可解释概念的稀疏概率组合。这些结果表明,CLIP 潜在空间更忠实地被表征为超球面语义混合而非各向同性高斯,为建模和理解多模态表示建立了一个简单且几何一致的概率框架。
英文摘要
Contrastive Language-Image Pretraining (CLIP) representations form a semantic embedding space governed by cosine similarity, reflecting an intrinsic hyperspherical geometry. However, existing probabilistic interpretations typically rely on Gaussian assumptions, which fail to capture this directional and multimodal structure. We propose a principled density model for the CLIP latent space based on Mixtures of von Mises-Fisher (MovMF) distributions defined on the unit hypersphere. Using the Expectation-Maximization (EM) algorithm, we efficiently learn a probabilistic model in which each mixture component corresponds to a coherent semantic concept. This formulation yields a closed-form likelihood naturally aligned with hyperspherical geometry, enabling accurate and interpretable density estimation. Empirically, our model significantly improves long-tailed and out-of-distribution detection and provides a natural semantic decomposition, representing each embedding as a sparse probabilistic combination of interpretable concepts. These results suggest that CLIP latent space is more faithfully characterized as a hyperspherical semantic mixture rather than an isotropic Gaussian, establishing a simple and geometrically consistent probabilistic framework for modeling and understanding multimodal representations. Project page is available at https://xiaoyuzhizi.github.io/movmf-clip/.
Comments23 pages, 8 figures