视觉语言模型推理中的视觉访问边界
What Keeps Vision-Language Models Looking at the Image?
浏览论文内容
中文总结 AI 辅助
研究视觉语言模型推理中CoT提示扩展的内容,通过视觉访问扫描定义VAB,发现CoT不主要靠延长图像令牌访问提效,而是扩展语言侧计算,且CoT收益受感知读出限制,瓶颈在读出非计数。
中文摘要 AI 辅助
思维链(CoT)提示被广泛用作视觉语言模型(VLM)的测试时扩展策略,但尚不清楚VLM生成更长推理轨迹时扩展的是什么。我们探讨CoT是否需要持续访问图像令牌,还是主要在前向传播早期可用的视觉信息上运行。我们引入视觉访问扫描,一种因果干预,沿层深度和生成时间屏蔽从生成令牌查询到图像令牌键的注意力,并将视觉访问边界(VAB)定义为保持任务准确性的最小访问区域。在Qwen2.5-VL和InternVL3的六种模型配置中,无CoT直接回答和CoT提示都表现出有限的VAB。在Qwen2.5-VL-32B和InternVL3的14B和38B规模下,当将CoT与无CoT全访问目标进行评估时,其VAB层与无CoT边界最多相差两层,尽管生成时间长得多。这表明CoT不是通过在整个推理轨迹中延长直接图像令牌访问来提高性能,而是通过扩展对图像派生隐藏状态信息的语言侧计算。我们进一步表明CoT的收益受感知读出的限制。当模型可以可靠地读出查询的视觉属性时,CoT有帮助,但当读出不可靠时则不然。一个符号属性预言机表明,一旦将真实属性作为文本提供,CoT可以改善计数,而一个单对象探测与解码检查表明,硬属性可以从隐藏状态线性恢复,但模型本身难以输出。总之,这些分析将瓶颈置于读出而非计数。
英文摘要
When do vision-language models need direct access to the image while generating an answer? We study image dependence during answer generation by examining how the visual information needed for the current question becomes available in context. We intervene on direct image access while retaining previously computed states. Across real-image and synthetic tasks, we show that, depending on the generation process, direct access can continue to support accuracy after question processing. Supplying the required attributes as text in the context weakens this dependence. On synthetic tasks, we also examine how dependence changes as the model itself states the required attributes. Before attribute expression, severing access reduces accuracy, and replacing image-side states shifts answers toward the counterfactual content. After sufficient expression, both interventions have smaller effects. Even with an identical generated prefix, dependence differs according to whether image access was available during question processing. Thus, both the visible text and the preceding image access matter. Several of these patterns hold across model families, including Qwen2.5-VL-32B and InternVL3-14B. These findings offer a view of image dependence in terms of the information needed for the current question and the history of image access, beyond generation position alone. This perspective provides a basis for deciding when to reduce visual access during an answer and which visual information to retain for subsequent questions.
发表机构
- The University of Tokyo(东京大学)
机构由 AI 辅助整理,请以论文原文为准。