大型视觉-语言模型中的涌现目标导向注意力
Emergent Goal-Directed Attention in Large Vision-Language Models
浏览论文内容
中文总结 AI 辅助
研究发现,无需注视监督,通用视觉-语言模型(如Qwen3-VL-32B-Thinking和Gemma-4-26B-A4B-it)在视觉搜索和自由观看任务中能涌现与人类目标导向注意力对齐的空间优先级,为预测人类注视提供可扩展工具。
中文摘要 AI 辅助
人类观察者会根据任务目标优先处理视觉信息。大多数自然观看的计算模型都是针对自由观看进行注视训练的,这留下了一个问题:在没有注视监督的系统中,目标导向的注意力是否能够涌现。我们在4,887个自然场景中,在视觉搜索和自由观看指令下,测试了两个现成的视觉-语言模型(VLMs),即Qwen3-VL-32B-Thinking和Gemma-4-26B-A4B-it。模型预测与人类在相同图像上对应任务下的注视点进行了比较。两个模型在匹配目标下比在非匹配目标下更接近人类注视点。这种交叉效应在目标缺失场景中持续存在,在那里对齐无法用简单的视觉基础解释,并且出现在解码器层的读出中。此外,模型思考轨迹在搜索过程中基于目标语义,在自由观看中基于视觉显著性。这些发现表明,通用视觉-语言模型可以在没有特定注视训练的情况下产生与人类对齐的目标导向空间优先级,为目标导向注意力理论提供信息,并提供可扩展的工具来预测人们在各种任务中的注视位置。
英文摘要
Human observers prioritize visual information according to task goals. Most computational models of naturalistic viewing are gaze-trained for free viewing, leaving open whether goal-directed attention can emerge in systems without gaze supervision. We tested two off-the-shelf vision-language models (VLMs), Qwen3-VL-32B-Thinking and Gemma-4-26B-A4B-it, on 4,887 naturalistic scenes under visual-search and free-viewing instructions. Model predictions were compared with human fixations on the same images under corresponding tasks. Both models aligned more closely with human fixations under matching goals than under mismatched goals. This crossover persisted in target-absent scenes, where alignment could not be explained by simple visual grounding, and appeared in decoder-layer readouts. Furthermore, model-thinking traces were grounded in target semantics during search and in visual prominence during free viewing. These findings show that general-purpose VLMs can generate human-aligned, goal-directed spatial priorities without gaze-specific training, informing theories of goal-directed attention and offering scalable tools for predicting where people look across tasks.
发表机构
- University of Michigan Transportation Research Institute(密歇根大学交通研究所)
机构由 AI 辅助整理,请以论文原文为准。