看清实际存在的东西:用于在视觉语言模型中对代理视觉证据进行反事实评估的PriVE-Bench和PriVE-Tools
Seeing What Is Actually There: PriVE-Bench and PriVE-Tools for Counterfactual Evaluation of Agentic Visual Evidence in VLMs
浏览论文内容
中文总结 AI 辅助
本文针对视觉语言模型常依先验而非图像回答问题的现象,引入PriVE-Bench和PriVE-Tools。通过配对图像及工具衍生证据评估模型,对比不同输入,发现视觉证据工具在特定情况有帮助,但不能解决所有模型依先验回答的问题。
中文摘要 AI 辅助
视觉语言模型(VLMs)常依据学习到的语言和类别先验回答视觉问题,而非基于图像本身。反事实图像为这种失败模式提供了自然诊断场景。现有反事实基准主要询问先验跟随行为是否存在,本文进一步探讨:额外视觉证据视图能否帮助VLMs对抗先验?为此引入PriVE-Bench,用配对的原始图像和反事实图像区分基于视觉的答案和与先验一致的错误。还引入PriVE-Tools评估工具衍生的视觉证据是否改善反事实冲突下的基础。通过对比不同输入,结果显示视觉证据工具在某些情况下有帮助,但并非通用解决方案。
英文摘要
Vision-language models (VLMs) often answer visual questions using learned language and category priors rather than grounding their predictions in the image itself. Counterfactual images provide a natural diagnostic setting for this failure mode: when visible evidence contradicts what is usually true, a grounded model should answer from the pixels, while a prior-following model will produce a canonical but visually incorrect response. However, existing counterfactual benchmarks mainly ask whether such prior-following behavior exists. In this paper, we ask a further question motivated by the rise of tool-augmented and agentic vision systems: can additional visual evidence views help VLMs reason against their priors? We introduce PriVE-Bench, a Prior-vs-Visual Evidence Benchmark that uses paired original and counterfactual images to distinguish visually grounded answers from prior-consistent errors. We further introduce PriVE-Tools, a controlled agentic-vision-inspired extension that evaluates whether tool-derived visual evidence -- including bounding boxes, crops, zoom panels, and contours -- improves grounding under the same counterfactual conflicts. Across open- and closed-source VLMs, we compare raw, paired-image, and tool-conditioned inputs using accuracy, prior-following error rate, and other-response rate. Our results show that visual evidence tools can help in some settings, especially when models can use localized evidence effectively, but they are not a universal remedy: several models continue to follow language and category priors even when relevant visual evidence is explicitly provided.
发表机构
- The University of Manchester(曼彻斯特大学)
- The University of Melbourne(墨尔本大学)
- University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- University of Edinburgh(爱丁堡大学)
- University of Southern California(南加州大学)
- Fudan University(复旦大学)
- University of Newcastle(纽卡斯尔大学)
机构由 AI 辅助整理,请以论文原文为准。