arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

相似的选择,不同的注意力:人类与视觉-语言模型中的跨模态关联

Similar Choices, Different Attention: Cross-Modal Associations in Humans and Vision-Language Models

Sumin Hong, Katsumi Ibaraki, Renee Shi, David Chiang, Toby Jia-Jun Li

arXiv 2609.36475首次发表:更新:

发表机构

University of Notre Dame(圣母大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过相同刺激下的选择与眼动实验,发现大型视觉-语言模型虽在选择上对齐人类,但其注意力与人类注视匹配度低于中心偏置基线,表明匹配选择或注视不足以证明跨模态处理与人类对齐。

AI 中文摘要

跨模态关联是不同模态间特征的系统性配对,例如“bouba”与圆形、“kiki”与尖锐形状的关联。先前的工作在人类与视觉-语言模型(VLMs)之间比较了此类关联,但常常在人类与模型之间使用不同的刺激或任务。在此,我们探究VLMs是否不仅在选择上与人类对齐,而且在做出这些选择时注视的位置上也与人类对齐。我们研究了VLMs和人类(N=53),向他们呈现相同的刺激——一个伪词和两幅图像,并记录参与者的选择与眼动数据,这些数据我们予以公开。我们发现,在几个较大的VLMs中存在选择对齐,但其显著性图与人类注视的匹配程度低于中心偏置基线(即每幅图像中心的一个固定高斯分布)。在人类选择上对小型VLMs进行微调,使其选择对齐达到人类多数投票参考在未见词和图像上的水平,但其注意力与人类注视的匹配程度仍低于该基线。在人类注视数据上训练模型注意力,能提高注意力-注视相关性,但未改善选择对齐;而每幅图像位置上的单一平均注视图也能以相似幅度提高该相关性。因此,匹配人类选择甚至人类注视模式,并不足以证明存在与人类对齐的跨模态处理。

英文摘要

Cross-modal associations are systematic pairings of features across modalities, such as the association of 'bouba' with round shapes and 'kiki' with sharp shapes. Prior work has compared humans and vision-language models (VLMs) on such associations, but often using different stimuli or tasks between humans and models. Here, we ask whether VLMs align with humans not only in choices, but also in where they look when making those choices. We study both VLMs and humans (N = 53), presenting them with the same stimuli, a pseudo-word and two images, and record participants' choices and eye movements, which we release. We find choice alignment in a few larger VLMs, but their saliency matches human gaze less closely than a center-bias baseline, a fixed Gaussian at the center of each image. Fine-tuning small VLMs on human choices brings their choice alignment to the level of a human majority-vote reference on unseen words and images, yet their attention still matches human gaze less closely than this baseline. Training model attention on human gaze raises attention-gaze correlation without improving choice alignment, and a single average gaze map per image position raises it by a similar amount. Matching human choices, or even human gaze patterns, is therefore not sufficient evidence of human-aligned cross-modal processing.

Comments9 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑