交叉质询下的视觉证据:评估与控制视觉语言模型中的决策级证据使用
Visual Evidence Under Cross-Examination: Evaluating and Controlling Decision-Level Evidence Use in Vision-Language Models
浏览论文内容
中文总结 AI 辅助
本研究提出CROSS-Bench基准和RIVET接口,通过失效与重绑定测试评估视觉语言模型中的候选绑定证据使用,证明证据效用与效应归属可分离,并显著提升决策准确率。
中文摘要 AI 辅助
视觉语言模型越来越多地通过裁剪区域、局部区域和工具生成的观察结果进行推理。然而,一个观察结果可能影响答案,却对其所支持的候选对象没有益处。我们研究了候选绑定视觉贡献:有效证据应有助于候选对象,使其支持关系失效应移除其额外影响,而有效的重新绑定应将这种影响重定向到新支持的候选对象。我们引入了CROSS-Bench,一个包含28,000个决策问题的基准,并在一个专门的评估子集上配备了匹配的失效和重新绑定测试。我们的RIVET接口保留了证据的身份和不确定性,组合了候选条件响应,并分别控制其强度。共享证据实验表明,任务准确性和证据归属可能发生分歧。在匹配容量和训练条件下,RIVET将归一化效应转移从0.512提高到0.651,其中干净证据具有正面效应。该优势在常见评估示例和重复决策层拟合中持续存在。当从原始输入预测证据时,相对于没有辅助证据的相同模型,RIVET在四个冻结骨干网络上平均将CROSS-Bench准确率提高了5.70个百分点。这些结果将视觉证据的效用与其效应的候选特定目的地区分开来。
英文摘要
Vision-language models increasingly reason through crops, regions, and tool-produced observations. Yet an observation can influence the answer without benefiting the candidate it supports. We study candidate-bound visual contribution: valid evidence should help, invalidating its supporting relation should remove its additional effect, and valid rebinding should redirect that effect to the newly supported candidate. We introduce CROSS-Bench, a benchmark of 28,000 decision problems, with matched invalidation and rebinding tests on a dedicated evaluation subset. Our RIVET interface preserves evidence identity and uncertainty, composes a candidate-conditioned response, and separately controls its strength. Shared-evidence experiments show that task accuracy and evidence ownership can diverge. Under matched capacity and training, RIVET increases normalized effect transfer from 0.512 to 0.651 where clean evidence has a positive effect. The advantage persists on common evaluation examples and across repeated decision-layer fits. With evidence predicted from raw inputs, RIVET improves CROSS-Bench accuracy by an average of 5.70 pp across four frozen backbones, relative to the same models without auxiliary evidence. These results separate the utility of visual evidence from the candidate-specific destination of its effect.
发表机构
- Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences(中国科学院空间应用工程与技术中心)
- University of Chinese Academy of Sciences(中国科学院大学)
- University of Science and Technology of China(中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。