证据顺序校准:在问题关键证据渐进丢失下的选择性视觉推理
Evidence-Order Calibration for Selective Visual Reasoning under Progressive Loss of Question-Critical Evidence
浏览论文内容
中文总结 AI 辅助
本研究提出证据顺序校准方法,通过后置可靠性头监督证据单调性,在渐进掩蔽问题关键证据时降低违反率,区分顺序一致性与正确性判别。
中文摘要 AI 辅助
视觉语言模型(VLM)的置信度在视觉证据被降级时可能在总体上发生变化,而在单个示例内部却保持结构上的不一致。我们沿着五步、以问题为条件的证据丢失轨迹研究答案级别的可靠性。使用冻结的Qwen2.5-VL-3B-Instruct模型,我们通过逐步掩蔽场景图定位的问题关键区域,构建了176条可接受的GQA衍生轨迹(880种掩蔽条件)。原生序列置信度的证据单调性违反率(EMVR)为0.436,且92.0%的轨迹包含至少一个相邻违反。匹配的非关键区域对照显示,完全关键掩蔽使准确率降低28.2个百分点,而同等大小的非关键掩蔽仅降低0.6个百分点;配对差异为27.6个百分点(95%置信区间[20.0, 34.7])。我们在冻结的隐藏状态、序列置信度和熵上训练了一个轻量级的后置可靠性头。在二元交叉熵(BCE)中加入证据顺序监督,将掩蔽EMVR从0.330降至0.303(配对差异-0.027,95%置信区间[-0.044, -0.010])。相同的掩蔽训练目标在未见过的局部高斯模糊下,对保留的问题ID将EMVR从0.449降至0.402(差异-0.0468,95%置信区间[-0.0739, -0.0199])。两个学习头之间的AUROC、Brier和AURC差异在统计上尚无定论,而原生置信度在选择性风险排序方面仍然更强。结果将证据顺序一致性与传统的正确性判别区分开来,而非确立通用的置信度优越性。
英文摘要
Vision-language model (VLM) confidence may change in aggregate when visual evidence is degraded while remaining structurally inconsistent within individual examples. We study answer-level reliability along five-step, question-conditioned evidence-loss trajectories. Using a frozen Qwen2.5-VL-3B-Instruct model, we construct 176 accepted GQA-derived trajectories (880 masking conditions) by progressively masking scene-graph-localized question-critical regions. Native sequence confidence has an evidence monotonicity violation rate (EMVR) of 0.436, and 92.0% of trajectories contain at least one adjacent violation. A matched non-critical-region control shows that full critical masking reduces accuracy by 28.2 percentage points, compared with 0.6 points for equally sized non-critical masks; the paired difference is 27.6 points (95% CI [20.0, 34.7]). We train a lightweight post-hoc reliability head on frozen hidden states, sequence confidence, and entropy. Adding evidence-order supervision to binary cross-entropy (BCE) reduces masking EMVR from 0.330 to 0.303 (paired difference -0.027, 95% CI [-0.044, -0.010]). The same mask-trained objective reduces EMVR from 0.449 to 0.402 on held-out question IDs under unseen local Gaussian blur (difference -0.0468, 95% CI [-0.0739, -0.0199]). AUROC, Brier, and AURC differences between the two learned heads are statistically inconclusive, and native confidence remains stronger for selective-risk ranking. The results separate evidence-order consistency from conventional correctness discrimination rather than establishing generic confidence superiority.
发表机构
- University of Moratuwa(莫拉图瓦大学)
机构由 AI 辅助整理,请以论文原文为准。