图像何时决定答案?跨图表与场景的视觉可回答性基准测试
When Does an Image Determine the Answer? Benchmarking Visual Answerability across Charts and Scenes
浏览论文内容
中文总结 AI 辅助
该研究提出一个跨图表、场景和照片的视觉可回答性基准,要求模型在证据充分时回答、不足时弃权,并揭示仅评估答案决策会掩盖完整任务失败。
中文摘要 AI 辅助
可靠的视觉问答要求在证据充分时给出正确答案,在证据不足时弃权(不执行)。我们引入了一个基准测试,将完整问题评估与跨PlotQA图表、CLEVR渲染场景和GQA照片的显式证据标签联系起来。每个问题将原始图像和编辑后的图像分组,并独立呈现;成功要求每个支持的答案和每个必需的弃权(不执行)都正确。对于图表缺失信息标签,可执行见证者证明,可接受的完整图表在掩蔽后给出不同答案但像素相同。场景标签遵循源程序和编辑,并对照片进行残差线索分析。在来自六个模型配置的72,000个响应中,观察到的最高完整任务成功率分别为57.0%、43.5%和33.7%。在图表上,最强配置实现了96.2%的逐视图决策准确率,但其835个组中265个组的所有决策正确,仍包含不正确的答案。将支持的答案和必要的弃权(不执行)一起评估,暴露了仅靠可回答性决策所掩盖的失败。
英文摘要
Reliable visual question answering requires correct answers when evidence is sufficient and abstention when it is not. We introduce a benchmark that connects complete-question evaluation with explicit evidence for its labels across PlotQA charts, CLEVR rendered scenes, and GQA photographs. Each question groups original and edited images, presented independently; success requires every supported answer and every required abstention to be correct. For chart missing-information labels, executable witnesses establish that admissible complete charts give different answers but identical pixels after masking. Scene labels follow source programs and edits, with a residual-cue analysis for photographs. Across 72,000 responses from six model configurations, the highest observed complete task success rates are 57.0%, 43.5%, and 33.7%, respectively. On charts, the strongest configuration achieves 96.2% per-view decision accuracy, yet 265 of its 835 groups with every decision correct still contain incorrect answers. Evaluating supported answers and necessary abstentions together exposes failures that answerability decisions alone conceal.
发表机构
- LG Uplus(LG U+)
机构由 AI 辅助整理,请以论文原文为准。