arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10224cs.CV

UOT-Gap:基于非平衡最优传输的视觉-语言模型模态间隙变分原理

UOT-Gap: A Variational Principle for the Modality Gap in Vision-Language Models via Unbalanced Optimal Transport

  • Guangdong Police College(广东警官学院)

机构由 AI 辅助整理,请以论文原文为准。

Zonglin Yang, Huilan Ma, Xudan Zheng, Yuejun Xie

AI总结:

提出UOT-Gap,利用非平衡最优传输建模模态间隙,通过成对感知残差诊断标题质量与检索鲁棒性,在多个数据集上取得高相关性。

AI中文摘要:

CLIP等视觉-语言模型将图像和文本嵌入到共享空间中,但不同模态的分布往往保持分离。现有解释将这种模态间隙归因于初始化、对比学习动态和信息不平衡,但其对检索的分布性及成对贡献仍未解决。我们提出UOT-Gap,一种无需训练的变分诊断方法,利用非平衡熵最优传输(UOT)对冻结的图像和文本嵌入进行建模。UOT最优解分离了传输、耦合复杂度和边缘质量变化;一个互补的成对感知残差将观察到的图像-文本对与UOT软匹配进行比较。在Flickr8K和COCO-1K上使用冻结的CLIP、OpenCLIP和SigLIP编码器,标题退化使Flickr8K的Recall@1从0.559降至0.003。在六种数据集-模型条件下,成对感知残差跟踪检索退化,平均绝对Spearman相关系数为0.973,而平均间隙仅为0.392。该关联在五个随机COCO-1K子集上保持稳定,为0.954±0.026,最小值为0.943。UOT重心更新在降低传输目标的同时损害检索性能,从而区分了几何目标下降与任务改进。这些结果确立了UOT-Gap作为标题质量、模态对齐和检索鲁棒性的诊断工具。

英文摘要:

Vision-language models such as CLIP embed images and text in a shared space, where modality-specific distributions often remain separated. Existing accounts connect this modality gap to initialization, contrastive dynamics, and information imbalance, while its distributional and pairwise contributions to retrieval remain unresolved. We introduce UOT-Gap, a training-free variational diagnostic that models frozen image and text embeddings with unbalanced entropic optimal transport (UOT). The UOT optimum separates transport, coupling complexity, and marginal mass variation; a complementary pair-aware residual compares observed image-caption pairs with the UOT soft matching. On Flickr8K and COCO-1K with frozen CLIP, OpenCLIP, and SigLIP encoders, caption degradation reduces Flickr8K Recall@1 from 0.559 to 0.003. Across six dataset-model conditions, the pair-aware residual tracks retrieval degradation with mean absolute Spearman 0.973, compared with 0.392 for the mean gap. The association remains stable across five random COCO-1K subsets at $0.954\pm0.026$, with a minimum of 0.943. UOT barycentric updates reduce the transport objective while degrading retrieval, distinguishing geometric objective descent from task improvement. These results establish UOT-Gap as a diagnostic for caption quality, modality alignment, and retrieval robustness.

补充信息

↑