arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.23235cs.CV

一种基于重构的、超越参考字幕的字幕评估框架

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions

Zhijiang Tang, Jiaxin Qi, Kaihua Tang, Yuhua Zheng, Jianqiang Huang

首次发表
浏览论文内容

中文总结 AI 辅助

研究不依赖参考字幕评估图像字幕对语义的忠实度问题,提出基于重构的评估原则,通过测试重构在下游视觉-语言任务中与原图匹配度得无参考字幕分数,还刻画局限性并引入CTTD替代方案。

中文摘要 AI 辅助

图像字幕是视觉-语言研究中的一项主要任务,然而,在不依赖参考字幕的情况下,评估字幕对图像语义的忠实程度仍未得到解决。目前的评估依赖于人工标注的参考字幕,其内容反映了标注者的意图和字幕能力。在本文中,我们研究了一种基于重构的字幕评估原则:一个字幕的好坏取决于它能够重构原始图像的能力。然而,由于字幕本质上会压缩视觉信息,因此无法恢复所有细节,对重构图像和源图像进行逐像素比较既不可行也没有意义。通过我们对字幕本质的深入分析,其根本目的是传递图像的语义内容,我们提出了一个修订原则:一个字幕的好坏取决于它能够实现与原始图像语义等效的重构的能力。为了评估语义等效性,我们测试重构在一系列下游视觉-语言任务中是否与原始图像匹配,从而产生一个无参考、任务条件的字幕分数。我们描述了组件相关的局限性,并引入了成本更低的字幕图灵测试数据集(CTTD)替代方案。

英文摘要

Image captioning is a primary task in vision--language research, yet assessing how faithfully a caption preserves image semantics without relying on reference captions remains unsettled. Prevailing evaluations rely on human-annotated references, whose content reflects annotator intent and captioning proficiency. In this paper, we study a reconstruction-based principle for caption evaluation: a caption is as good as its capacity to enable reconstruction of the original image. However, because captioning inherently compresses visual information, it is impossible to recover all details, and pixel-wise comparison between reconstructed and source images is neither feasible nor meaningful. Through our in-depth analysis of the nature of captions, whose fundamental purpose is to transmit the semantic content of an image, we propose a revised principle: a caption is as good as its capacity to enable a reconstruction that is semantically equivalent to the original. To assess semantic equivalence, we test whether the reconstruction matches the original image across a suite of downstream vision--language tasks, yielding a reference-free, task-conditioned caption score. We characterize component-dependent limitations and introduce the lower-cost Captioning Turing Test Dataset (CTTD) surrogate.

发表机构

  • Computer Network Information Center, Chinese Academy of Sciences(中国科学院计算机网络信息中心)
  • Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences(中国科学院大学杭州高等研究院)
  • Tongji University(同济大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑