发表机构
Sword Health; NOVALINCS, NOVA University of Lisbon(Sword Health; NOVA里斯本大学NOVALINCS)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DEPICT提出一种免训练的图像-文本对齐评分指标,通过图像与字幕答案一致性替代固定参考,结合整体评分,显著提升否定检测准确率并超越现有方法。
AI 中文摘要
图像-文本对齐是计算机视觉中的一个核心问题,其应用包括字幕评估、幻觉检测、数据整理以及文本到图像(T2I)生成器的基准测试。随着T2I模型的改进,基准测试变得要求更高,需要能够发现一系列问题的指标,例如缺失对象、属性交换、计数错误和忽略的否定。近期的工作通过偏好数据微调评估器,或通过提示视觉语言模型(整体使用字幕或使用分解的验证问题)来解决这一问题。然而,现有方法存在不足:微调指标仍受限于单一骨干网络和训练分布;整体指标遗漏细粒度细节;分解指标依赖于固定的“是”假设,当该假设失败时会惩罚忠实图像。相比之下,我们提出DEPICT,一种免训练指标,用基于图像和仅字幕答案之间的预期一致性取代固定参考答案,并根据字幕决定问题的决定性程度对问题加权。通过替换固定参考,我们的一致性规则将否定准确率从19%提升至88%。为恢复分解过程中丢失的上下文,DEPICT将此一致性得分与整体得分合并。我们在五个基准和来自三个模型家族的十一个骨干网络上评估DEPICT,发现它超越了所有免训练指标,并在三个人类相关性基准中的两个上超过了微调评估器。
英文摘要
Image-text alignment is a core problem in computer vision with applications in caption evaluation, hallucination detection, data curation, and the benchmarking of text-to-image (T2I) generators. As T2I models improve, benchmarking has become demanding, requiring metrics capable of finding a series of issues like missing objects, swapped attributes, miscounts, and ignored negations. Recent work addresses this by fine-tuning evaluators on preference data or by prompting a vision-language model, either holistically with the caption or with decomposed verification questions. However, existing approaches fall short: fine-tuned metrics remain bound to one backbone and training distribution; holistic metrics miss fine-grained details; and decomposed metrics rely on a fixed-YES assumption that penalizes faithful images whenever that assumption fails. In contrast, we propose DEPICT, a training-free metric that replaces fixed reference answers with expected agreement between image-based and caption-only answers, weighting questions by how decisively the caption determines them. By replacing fixed references, our agreement rule increases negation accuracy from 19% to 88%. To recover the context lost during decomposition, DEPICT merges this agreement score with a holistic score. We evaluate DEPICT on five benchmarks and eleven backbones from three model families and find that it surpasses all training-free metrics and exceeds fine-tuned evaluators on two out of three human-correlation benchmarks.