当模态差距缩小失效时:CLIP中的预测层面中心性
When Modality Gap Reduction Fails: Prediction-Level Hubness in CLIP
浏览论文内容
中文总结 AI 辅助
本文针对CLIP中模态差距缩小无法持续提升零样本准确率的问题,提出预测层面中心性这一失效模式,表明差距校正会导致预测集中,需结合下游预测结构评估校正效果。
中文摘要 AI 辅助
缩小CLIP中图像与文本表征之间的模态差距被广泛认为可提升跨模态对齐效果及下游任务性能。然而,更小的平均图像-文本差距并不一定能带来持续的准确率提升。本文从零样本分类的决策结构视角分析这一不匹配问题,即针对输入图像选择最相似的类别文本原型。零样本准确率不仅取决于平均图像-文本对齐程度,还取决于按类别划分的决策边界。以线性校正作为解析上可处理的案例,本文表明模态差距校正会改变各类别间的相对决策结构,导致预测结果集中于少数类别。本文将这种输出空间中的失效模式称为预测层面中心性。此外,在多个数据集上开展的实验显示,对于线性校正及基于学习的校正方法,差距校正下的准确率下降始终与预测集中度提升相关联。这从下游决策结构的角度为模态差距缩小无法持续提升CLIP零样本准确率提供了系统性解释。研究结果表明,对差距校正的评估不应仅基于平均对齐效果,还需考虑其对下游预测结构的影响。
英文摘要
Reducing the modality gap between image and text representations in CLIP is widely expected to improve cross-modal alignment and downstream performance. However, a smaller average image-text gap does not necessarily lead to consistent accuracy gains. We analyze this mismatch from the perspective of the decision structure in zero-shot classification, i.e. selecting the most similar class-text prototype for an input image. Zero-shot accuracy depends not only on average image--text alignment, but also on class-wise decision margins. Using Linear correction as an analytically tractable case, we show that modality gap correction can alter the relative decision structure among classes and cause predictions to concentrate on a small subset of classes. We refer to this output-space failure mode as prediction-level hubness. Furthermore, experiments across multiple datasets show that accuracy degradation under gap correction is consistently associated with increased prediction concentration, both for Linear correction and for learning-based correction methods. This provides a systematic explanation of why modality gap reduction does not consistently improve CLIP zero-shot accuracy from the perspective of downstream decision structure. Our results suggest that gap correction should be evaluated not only by average alignment, but also by its impact on downstream prediction structure.
发表机构
- Hitotsubashi University(一桥大学)
机构由 AI 辅助整理,请以论文原文为准。