发表机构
Indian Institute of Technology Jodhpur(印度理工学院焦特布尔分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过几何分析揭示视觉-语言模型中模态间隙的秩一结构,统一解释了间隙修改在不同下游任务中的改善、降低或恢复性能的效应,并为间隙干预选择提供原则性指导。
AI 中文摘要
对比式视觉-语言模型通过对齐匹配的图像-文本对来学习共享嵌入空间,但其表示仍被模态间隙所分隔。先前的研究报告了修改这一间隙的不同效应:减小间隙可以改善零样本分类和跨模态对齐,而移除间隙相关结构则可能降低图像-文本检索性能。在本文中,我们为这些任务依赖的效应提供了统一的几何解释。在CLIP和SigLIP编码器中,我们发现一个单一主导方向捕获了图像-文本均值分离平方范数的94.4%-99.9%,揭示出均值分离分量近似为秩一。对相似度分数的分解随后识别出三个任务特定的角色。在零样本分类中,查询侧固定间隙偏移减法恰好等价于加性类别偏置。在标准跨模态检索中,投影出间隙方向并重新归一化残差会丢弃候选特定的范数信息,导致乘性排序失真;一个几何导出的指数追踪了网格搜索最优值(Spearman rho = 0.93),并在某些设置中恢复了性能,尽管增益的转移并不均匀。在混合模态检索中,间隙方向按模态对候选进行排序;移除它可以改善跨模态排序,这与随机或非间隙对照不同。移除后的残余语义结构定义了秩一解释的极限。这些结果共同解释了为什么间隙修改可以在不同下游设置中改善、降低或恢复性能。通过阐明间隙修改何时以及为何改变模型行为,这一解释为在评估的下游任务中基于相似性的视觉-语言系统选择间隙干预提供了原则性基础。
英文摘要
Contrastive vision-language models learn shared embedding spaces by aligning matched image-text pairs, yet their representations remain separated by a modality gap. Prior work reports divergent effects of modifying this gap: reducing it can improve zero-shot classification and cross-modal alignment, whereas removing gap-related structure can degrade image-text retrieval. In this paper, we provide a unified geometric explanation for these task-dependent effects. Across CLIP and SigLIP encoders, we find that a single dominant direction captures 94.4-99.9% of the squared norm of the image-text mean separation, revealing that the mean-separation component is approximately rank-one. A decomposition of the similarity score then identifies three task-specific roles. In zero-shot classification, query-side fixed gap-offset subtraction is exactly equivalent to an additive class bias. In standard cross-modal retrieval, projecting out the gap direction and renormalising residuals discards candidate-specific norm information, inducing a multiplicative ranking distortion; a geometry-derived exponent tracks the grid-search optimum (Spearman rho = 0.93) and restores performance in some settings, although the gains transfer unevenly. In mixed-modal retrieval, the gap direction sorts candidates by modality; its removal can improve cross-modal ranking, unlike random or non-gap controls. Residual semantic structure after removal defines the limits of the rank-one account. Together, these results explain why gap modification can improve, degrade, or restore performance across downstream settings. By clarifying when and why gap modification changes model behavior, this account provides a principled basis for selecting gap interventions in similarity-based vision-language systems across evaluated downstream tasks.
Comments9 pages, 3 figures, 3 tables