发表机构
Beijing University of Posts and Telecommunications(北京邮电大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究简答题VQA基准中模型答案语义正确性与和评估者期望表面形式匹配度被混淆的问题,通过人工验证语义判断等方法研究六个视觉语言模型和六个基准,发现答案类型影响不稳定性,提出官方分数应伴有语义审计和答案类型诊断以保持可解释性。
AI 中文摘要
简答题VQA基准将两个不同的量混为一谈:模型答案在语义上是否正确,以及该答案是否与自动评估者期望的表面形式匹配。我们使用人工验证的语义判断(精度为97.6%)对六个视觉语言模型和六个基准进行了研究,审计了超过37000个官方错误。另一个纯文本判断重现了相同的基准级假阴性模式,表明这种影响不是单个审计模型的产物。在文本丰富的基准上,这些错误中多达一半是语义上可接受的答案,纯粹因表面形式不匹配而受到惩罚。这种不稳定性由答案类型构成:提取式和多跨度答案比标量答案对评估者更敏感。良性提示和上下文重写进一步破坏了官方结果的稳定性,在不改变底层任务的情况下,以相当高的比率翻转项目级的正确性。确定性的仅CPU合约修复证实了漏计部分是可恢复的。这些发现意味着官方简答题VQA分数应伴有语义审计和答案类型诊断,以便保持可解释性。
英文摘要
Short-answer VQA benchmarks conflate two distinct quantities: whether a model's answer is semantically correct, and whether that answer matches the surface form expected by the automatic evaluator. We study this conflation across six vision--language models and six benchmarks, using a human-validated semantic judge (97.6% precision) to audit over 37k official errors. A second text-only judge reproduces the same benchmark-level false-negative pattern, showing that the effect is not an artifact of a single audit model. On text-rich benchmarks, up to half of these errors are semantically acceptable answers penalized purely for surface-form mismatch. This instability is structured by answer type: extractive and multi-span answers are far more evaluator-sensitive than scalar answers. Benign prompt and context rewrites further destabilize official outcomes, flipping item-level correctness at substantial rates without changing the underlying task. A deterministic CPU-only contract repair confirms that the undercount is partially recoverable. These findings imply that official short-answer VQA scores should be accompanied by semantic audits and answer-type diagnostics to remain interpretable.