发表机构
University of Chinese Academy of Sciences; Institute of Automation, Chinese Academy of Sciences(中国科学院大学; 中国科学院自动化研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出解耦式标题评估基准CAPEval,将标题质量分为覆盖度与精确度,发现二者分别对应理解与生成任务性能,为标题生成器的选择优化提供指导。
AI 中文摘要
标题是多模态理解和文本到图像生成的主要监督信号,但以往的评估将标题质量视为单一标量目标,混淆了两个不同属性:一是标题覆盖的视觉信息量,二是图像对其表述主张的可靠程度。为此,我们设计了解耦式标题评估基准CAPEval(Coverage And Precision Evaluation,覆盖度与精确度评估),包含人工撰写的真实标题和人工验证的原子清单条目。具体而言,CAPEval将标题质量分解为覆盖度和精确度:前者量化标题覆盖真实事实内容的全面程度,后者反映标题中所有主张的事实正确率。我们选取10个标题生成器,并以标题来源为唯一变量,在四个模型家族上开展受控的下游端到端实验。实证发现存在一致的任务依赖型解离:覆盖度是理解性能的更强关联因素,而精确度是生成性能的主导预测因子。这种解耦评估范式不仅能对标题质量进行更细粒度的诊断,还能为针对不同下游任务选择和优化标题生成器提供可操作的指导。
英文摘要
Captions serve as a primary supervision signal for both multimodal understanding and text-to-image generation. However, previous evaluations treat the caption quality as a single scalar objective, which conflates two distinct properties: (1) how much visual information a caption covers and (2) how reliably the image supports its stated claims. To this end, we design a decoupled caption evaluation benchmark, CAPEval (Coverage And Precision Evaluation), with human-written ground-truth captions and human-verified atomic checklist items. Specifically, CAPEval decomposes caption quality into Coverage and Precision. The former quantifies how thoroughly a caption covers ground-truth factual content, while the latter reflects the factual correctness rate of all claims expressed in the caption. We select 10 captioners and further conduct controlled downstream end-to-end experiments with them from four model families, where the caption source is the only variable. Empirically, we find a consistent task-dependent dissociation: Coverage serves as the stronger correlate for understanding performance, whereas Precision acts as the dominant predictor for generation performance. This decoupled evaluation paradigm not only delivers a more fine-grained diagnosis of caption quality, but also offers actionable guidance for selecting and optimizing captioners tailored to different downstream tasks.
Comments21 pages, 8 figures. Code and dataset will be available at https://liuzhipenggg.github.io/CAPEval/