不要只看,要干预:基于扰动的VQA图像区域标注
Don't Just Look, Intervene: Perturbation Based Region Labeling for VQA Images
- Arizona State University(亚利桑那州立大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出基于反事实干预的CSGR区域标注方法,为VQA图像生成因果视觉证据,并在多种接地感知训练中验证其提升效果。
AI中文摘要:
视觉语言模型(VLMs)应依赖于直接决定正确答案的视觉证据,但用于引导视觉推理的监督信息通常需要人工获取,成本高昂,或受限于特定数据集的注释原语。我们转而引入模型因果视觉证据作为注释目标,将其定义为:对于给定的图像-问题对,其反事实干预会改变模型答案分布的一组图像区域。基于这一原则,我们提出了反事实搜索接地区域(CSGR)。CSGR是一个可扩展的流水线,它提出候选区域,对其进行扰动,衡量它们对答案敏感性的影响,并跨多个评判者聚合这些证据,以近似VQA数据中答案关键区域。为了评估CSGR注释是否包含有用的监督信号,我们将其插入三种现有的接地感知训练流程:注意力引导、视觉CoT微调和潜在视觉推理。这些实验测试了所提出的注释方案能否在多种区域标签消费方式中提供有用的监督信号,而不是引入一种新的使用方式。在与竞争性自动区域标注机制的对比中,CSGR注释在域内和域外评估中均提供了相对于仅交叉熵微调的最一致的改进,表明所提出的标注方案捕获了有用的区域级信息。
英文摘要:
Vision Language Models (VLMs) should rely on visual evidence that directly determines the correct answer, but supervision for grounding visual reasoning is often expensive to obtain manually or tied to dataset-specific annotation primitives. We instead introduce model-causal visual evidence as an annotation target, defined as the set of image regions whose counterfactual intervention changes a model's answer distribution for a given image-question pair. Based on this principle, we introduce Counterfactual Search for Grounding Regions (CSGR). CSGR is a scalable pipeline that proposes candidate regions, perturbs them, measures their effect on answer sensitivity, and aggregates this evidence across multiple judges to approximate answer-critical regions in VQA data. To assess whether CSGR annotations contain a useful supervision signal, we plug them into three existing grounding-aware training routines: attention steering, Visual CoTfinetuning, and latent visual reasoning. These experiments test whether the proposed annotation scheme can provide a useful supervision signal across multiple ways of consuming region labels, rather than introducing a new way of using them. Across competing automatic region-labeling mechanisms, CSGR annotations provide the most consistent gains over Cross Entropy-only finetuning in both in-domain and out-of-domain evaluations, indicating that the proposed labeling scheme captures useful region-level information.