看关键之处:面向细粒度视觉-语言推理的免费、无标签视觉证据信号
Look Where It Counts: A Free, Label-Free Visual Evidence Signal for Fine-Grained Vision-Language Reasoning
- KGraph AI Solutions Pvt. Ltd.(KGraph AI解决方案私人有限公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出一种免费、无标签的视觉证据信号(对比证据差距),通过选择使模型答案分布最尖锐的裁剪区域来定位细粒度视觉证据,无需训练即可将准确率从70%提升至85%,但将其转化为训练方法仍待解决。
AI中文摘要:
多模态大语言模型(MLLMs)在细粒度视觉问题上失败,与其说是由于无法推理,不如说是由于从未看到证据:高分辨率图像在编码前被下采样,因此模型依据语言先验作答。标准的补救措施代价高昂:带标注的答案(SFT)、手工设计的验证器(RLVR)或大型外部教师模型(在线策略蒸馏)。我们探究视觉证据本身能否免费提供该信号。我们形式化了对比证据差距,即模型在给定与问题相关区域与无关区域条件下,对其自身输出的逐词元对数似然比,并在Qwen2.5-VL-7B、Qwen3-VL-8B和Qwen3-VL-30B-A3B上于V*Bench基准进行研究。我们的主要正面结果无需训练:使用单视图、无标签准则,选择使模型答案分布最尖锐的候选裁剪区域,即可发现包含答案的区域,无需边界框、训练或标签。该方法定位目标的性能比随机好4.4至5.1倍,并在推理时将细粒度准确率从70%提升至85%。我们进一步表明,该差距与模型自身的置信度互补。将两者结合预测正确性的性能优于单独使用任一者,AUC高达0.99,并在高置信度子集中标记自信但错误的答案,AUC范围为0.97至1.00。所有效应集中于感知瓶颈问题,并在全局上下文对照中消失。最后,我们报告一个诚实的负面结果:将同一信号转化为训练方法,即门控自蒸馏(SEG-Distill),在三种门控设计下于试点规模未能超越基础模型,而更激进的门控会降低准确率。该信号是真实的,但将其转化为训练收益仍是一个开放问题。
英文摘要:
Multimodal large language models (MLLMs) fail at fine-grained visual questions less because they cannot reason than because they never see the evidence: high-resolution images are downsampled before encoding, so the model answers from linguistic priors. The standard remedies are expensive: annotated answers (SFT), hand-engineered verifiers (RLVR), or a large external teacher (on-policy distillation). We ask whether the visual evidence itself can supply the signal for free. We formalize the contrastive evidence gap, the per-token log-likelihood ratio that a model assigns to its own output when conditioned on a question-relevant region versus an irrelevant one, and study it across Qwen2.5-VL-7B, Qwen3-VL-8B, and Qwen3-VL-30B-A3B on V*Bench. Our main positive result is training-free: selecting the candidate crop under which the model's answer distribution is most peaked, using a single-view, label-free criterion, discovers the answer-bearing region with no bounding boxes, training, or labels. It localizes the target 4.4 to 5.1 times better than chance and raises fine-grained accuracy from 70 percent to 85 percent at inference. We further show that the gap is complementary to the model's own confidence. Combining them predicts correctness better than either alone, with AUC up to 0.99, and flags confidently wrong answers, with AUC ranging from 0.97 to 1.00 within the high-confidence subset. All effects concentrate on perception-bottleneck questions and vanish on a global-context control. Finally, we report an honest negative result: converting the same signal into a training method, gated self-distillation (SEG-Distill), does not outperform the base model at pilot scale across three gate designs, while more aggressive gating degrades accuracy. The signal is real, but converting it into training gains remains an open problem.