arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CapProbe:通过全场景密集问答评估详细图像描述

CapProbe: Evaluating Detailed Image Captions via Full-Scene Dense Question Answering

Mouxiao Huang, Qiangyu Yan, Borui Jiang, Han Shu

arXiv 2608.11074首次发表:更新:

发表机构

Huawei Technologies(华为技术有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CapProbe是用于评估VLMs生成的详细图像描述的全场景密集问答基准,通过区域对齐的事实核查,揭示了13个VLMs的覆盖差距、能力-效率权衡及遗漏的失败模式,相关资源将很快发布。

AI 中文摘要

评估视觉语言模型(VLMs)生成的详细图像描述,不能仅停留在表层语义相似性层面。基于参考的指标(如CIDEr和SPICE)以及“LLM作为评分者”协议难以验证密集事实断言,而现有的基于问答(QA)的替代方案通常存在探测密度较低、领域覆盖较窄,或未明确建立单个问题与分割图像区域之间对齐关系的问题。我们提出CapProbe,这是一个全场景密集问答基准,将详细图像描述评估转化为区域对齐的事实核查任务。每张图像被分解为覆盖前景和背景元素的粗粒度语义区域;对于每个保留的区域,我们生成涵盖10个语义类别的多项选择题,形成密集的已探测视觉事实清单。在由37个L1领域和219个L2子领域组成的两级分类法指导下,CapProbe包含346张图像、1868个区域和25650个问题,平均每张图像对应74个问答对。语言评分者仅根据描述进行回答;“不确定”选项和有效准确率提供了一种依赖于评分者的代理指标,用于区分未回答的探测项与错误解决的探测项,而基于密度的指标会惩罚冗长但无信息的描述。该协议具有成本效益:通过将无约束的标量评分转化为结构化的多项选择题阅读任务,它减少了开放式评分偏差,同时保持了评分者条件性,并且在固定阅读器下产生相对稳定的模型排名。对13个VLMs的实验显示,模型之间存在较大的覆盖差距、明显的能力-效率权衡,以及稀疏或基于重叠的评估通常会遗漏的失败模式。基准数据、注释和评估代码将很快发布。

英文摘要

Evaluating detailed image captions from Vision-Language Models (VLMs) requires going beyond surface-level semantic similarity. Reference-based metrics (e.g., CIDEr and SPICE) and LLM-as-scorer protocols struggle to verify dense factual claims, while existing QA-based alternatives generally offer lower probe density, narrower domain coverage, or no explicit alignment between individual questions and segmented image regions. We introduce CapProbe, a full-scene dense QA benchmark that turns detailed caption evaluation into region-aligned factual checking. Each image is decomposed into coarse semantic regions covering both foreground and background elements; for every retained region, we generate multiple-choice questions spanning 10 semantic categories, forming a dense checklist of probed visual facts. Guided by a two-tier taxonomy of 37 L1 domains and 219 L2 sub-domains, CapProbe comprises 346 images, 1,868 regions, and 25,650 questions, averaging 74 QA pairs per image. A language judge answers from the caption alone; an Uncertain option and Effective Accuracy provide a judge-dependent proxy for distinguishing unanswered probes from incorrectly resolved ones, while density-based metrics penalize verbose yet uninformative captions. The protocol is cost-effective: by converting unconstrained scalar scoring into structured MCQ reading, it reduces open-ended scoring bias while remaining judge-conditioned and yields relatively stable model rankings under a fixed reader. Experiments on 13 VLMs show large Coverage gaps across models, a clear competency-efficiency trade-off, and failure modes that sparse or overlap-based evaluation often misses. The benchmark data, annotations, and evaluation code will be released soon.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑