arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.16214cs.CVcs.AI

是什么让语言表征成为人类大脑中高级视觉感知的良好模型?

What Makes Linguistic Representations Good Models of High-Level Visual Perception in the Human Brain?

Anna Bavaresco, Ina Klarić, Raquel Fernández, Marie-Francine Moens

首次发表
浏览论文内容

中文总结 AI 辅助

研究语言模型图像描述对人类大脑高级视觉感知的预测性,考虑六种字幕类型和五种语言模型,发现机器生成字幕表现优,文本嵌入器优于自回归语言模型,大脑预测性和行为一致性在中间网络深度达峰值,确立字幕嵌入为研究工具。

中文摘要 AI 辅助

用语言模型(LMs)表示的图像描述可预测人类大脑在高级视觉区域对自然图像的反应,但驱动这种预测性的因素尚不清楚。为了研究这一点,我们系统地研究了图像如何被描述以及使用哪些语言模型来嵌入这些描述。对于一组常见图像,我们考虑了六种字幕类型,包括人工标注和多种机器生成的字幕,它们在几个维度上有所不同。每个字幕用五种语言模型表示,涵盖用于预测后续单词的自回归语言模型和文本嵌入器。机器生成的字幕产生了显著的大脑预测性和一致性,通常超过了先前工作中使用的人工标注字幕。在所有字幕类型中,文本嵌入器始终优于自回归语言模型,在测量与图像相似性判断的行为一致性时也出现了这种模式。对来自不同模型层的字幕表征的分析进一步表明,大脑预测性和行为一致性在中间网络深度达到峰值,这一峰值出现在被认为标志着句法和语义结构出现的点之后不久。我们的结果表明,图像字幕的内容和用于表示它们的语言模型都会影响大脑和行为建模性能,将字幕嵌入确立为研究高级视觉感知的有用工具。

英文摘要

Image descriptions represented with language models (LMs) predict human brain responses to naturalistic images in high-level visual regions, but the factors driving this predictivity remain unclear. To investigate this, we systematically studied how images are described and which language models are used to embed those descriptions. For a common set of images, we considered six caption types -- including human-annotated and multiple machine-generated captions -- differing along several dimensions. Each caption was represented with five LMs, spanning autoregressive LMs trained to predict upcoming words and text embedders, i.e., LMs fine-tuned on semantic tasks requiring sentence/document-level representations. Machine-generated captions yielded significant brain predictivity and alignment, often surpassing human-annotated captions used in previous work. Across caption types, text embedders consistently outperformed autoregressive LMs, a pattern replicated when measuring behavioural alignment with image-similarity judgments. Analyses of caption representations from different model layers further revealed that both brain predictivity and behavioural alignment peak at intermediate network depth, shortly after a point thought to mark the emergence of syntactic and semantic structure. Altogether, our results demonstrate that both the content of image captions and the LM used to represent them influence brain- and behaviour-modelling performance, establishing caption embeddings as a useful tool for studying high-level visual perception.

发表机构

  • European Commission, Joint Research Centre (JRC)(欧盟委员会,联合研究中心 (JRC))

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑