发表机构
University of Stuttgart(斯图加特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出结合CLIP双编码器与双模态归因方法的新方法,无需微调或生成式架构,即可复现经典视觉世界研究的人类注视行为预测结果。
AI 中文摘要
神经语言模型的近期进展也推动了计算心理语言学领域的诸多研究,探讨神经语言模型是否也有望成为人类语言处理的模型。然而,相关研究大多聚焦于书面或口语的单模态场景。相比之下,视觉世界研究这类同时向参与者呈现视觉与语言输入的多模态实验范式却被忽视了。本文提出一种预测视觉世界研究中注视行为的新方法,该方法将CLIP系列的简单多模态双编码器模型与双模态归因方法相结合。我们证明了该方法能够可靠复现一项经典英语视觉世界研究的结果,该研究展示了人类的预测处理能力。值得注意的是,该方法无需生成式架构,也无需微调,尽管未针对此任务进行训练,却能实现上述效果。
英文摘要
The recent advances in neural language models have also spurred much work in computational psycholinguistics, asking whether neural LMs are also promising models of human language processing. However, work has been overwhelmingly focused on the unimodal case of written or spoken language. In contrast, multimodal experimental paradigms, like visual world studies that present participants with both visual and linguistic input simultaneously, have been neglected. In this paper, we present a novel approach that predicts gaze behavior in visual world studies. It does so by combining a simple multi-modal bi-encoder model of the CLIP family with a bimodal attribution method. We demonstrate the ability of this approach to robustly replicate the results of a seminal English visual world study which shows hu- man predictive processing. Remarkably, it does so without a generative architecture and without the need for fine-tuning, despite not being trained for this task.