AI 中文总结
该研究针对VLM自反思中存在的文本捷径问题,通过反事实分析诊断其特性,提出无需训练的FSAF干预,显著提升视觉更新率、降低先前答案率,为缓解VLM的文本复用偏差提供了有效方案。
AI 中文摘要
视觉语言模型(VLM)应在视觉证据变化时修正推理,此前未能做到这一点常归因于视觉注意力不足或上下文惯性,但模型复用了什么而非从当前图像重新计算仍不明确。我们发现,先前思维链(CoT)中承载证据的推理可形成文本捷径,在行为上与视觉重新计算竞争。对16个VLM的匹配反事实分析表明,承载证据的内容是先前CoT影响最稳健的载体;移除该承载证据的内容,比移除长度匹配的非证据上下文或最终答案跨度更能改变答案偏好,且随着更多过时证据被移除,先前控制会逐渐减弱。重新排序该证据也会削弱先前控制,显示其组织方式会调节捷径强度。除了直接答案,捷径在答案修正后仍会保留残余影响:削弱当前图像的支持会使偏好转回先前答案,而重复先前答案和复用前提主要在捷径仍活跃时出现。为限制这种影响,我们引入Fresh-State Attention Firewall(FSAF),一种无需训练的干预方法,可将新鲜计算与先前CoT隔离。在5个VLM上,FSAF将视觉更新率从35.28%提升至53.61%,并将先前答案率从39.22%降至3.67%。因此,可靠的VLM自反思不止需要再次查看:必须保护新鲜视觉重新计算免受过时文本复用的影响。
英文摘要
Vision-language models (VLMs) are expected to revise their reasoning when visual evidence changes. Failures to do so are often attributed to insufficient visual attention or contextual inertia, leaving unclear what models reuse instead of recomputing from the current image. We show that evidence-bearing reasoning in a prior chain of thought (CoT) can form a textual shortcut that competes behaviorally with visual recomputation. Across 16 VLMs, a matched counterfactual analysis identifies evidence-bearing content as the most robust carrier of prior-CoT influence. Removing this evidence-bearing content shifts answer preference more than removing length-matched non-evidence context or the final-answer span, with prior control weakening progressively as more stale evidence is removed. Reordering this evidence also weakens prior control, showing that its organization modulates shortcut strength. Beyond the immediate answer, the shortcut can retain residual influence after answer correction: weakening current-image support shifts preference back toward the prior answer, while repeated prior answers and reused premises arise mainly when the shortcut remains active. To limit this influence, we introduce Fresh-State Attention Firewall (FSAF), a training-free intervention that isolates fresh computation from the prior CoT. Across five VLMs, FSAF raises visual update rate from 35.28% to 53.61% and reduces prior-answer rate from 39.22% to 3.67%. Reliable VLM self-reflection therefore requires more than looking again: fresh visual recomputation must be protected from stale textual reuse.
Commentspreprint