长文本图文一致性评估中基于人类标注的校准方法
Human-Grounded Calibration for Long-Text Image-Text Congruence in Vision-Language Models
浏览论文内容
中文总结 AI 辅助
本文提出轻量级校准层CS,将长文本图像-文本一致性评分建模为校准的分数估计问题,并揭示检索性能、人类关联与阈值校准间的权衡。
中文摘要 AI 辅助
长文本图像-文本一致性评分对于必须评估详细文本描述是否与视觉内容匹配的视觉语言系统日益重要。然而,来自双编码器模型的原始相似度分数难以解释为校准的一致性度量,尤其是在图像和文本嵌入之间存在模态差距的情况下。本文提出了一致性分数(CS),一种轻量级校准层,将图像-文本相似性证据映射为有界分数。使用DOCCI和Urban1k数据集,我们评估了四个冻结的视觉语言骨干模型,并表明观察到的投影后质心距离的减少并不能统一提升图像-文本检索性能。基于人类标注的DOCCI评估进一步揭示了一种权衡:直接的后期校准保持了与人类判断的高度关联,而选定的基于投影的配置可以减少阈值相关的斜率和截距失真,但以牺牲检索性能和关联强度为代价。这些结果将长文本图像-文本一致性评分确立为一个校准的分数估计问题,其中检索性能、人类关联和阈值校准必须作为不同的目标进行评估。CS提供了一种轻量级的方式来揭示并操作这种分离。
英文摘要
Long-text image--text congruence scoring is increasingly important for vision-language systems that must evaluate whether detailed textual descriptions match visual content. However, raw similarity scores from dual-encoder models are difficult to interpret as calibrated congruence measures, especially under the modality gap between image and text embeddings. This paper proposes Congruency Score (CS), a lightweight calibration layer that maps image--text similarity evidence into a bounded score. Using DOCCI and Urban1k, we evaluate four frozen vision-language backbones and show that observed reductions in post-projection centroid distance do not uniformly improve image--text retrieval performance. Human-grounded evaluations on DOCCI further reveal a trade-off: direct post-hoc calibration preserves high association with human judgments, whereas selected projection-based configurations can reduce threshold-relevant slope and intercept distortions at the cost of retrieval performance and association strength. These results establish long-text image--text congruence scoring as a calibrated score-estimation problem, where retrieval performance, human association, and threshold calibration must be evaluated as distinct objectives. CS provides a lightweight way to expose and operationalize this separation.