arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

类人注意力?人类与多模态大语言模型视觉搜索的心理物理学比较

Human-Like Attention? A Psychophysical Comparison of Visual Search in Humans and MLLMs

Renchi Zhang, Joost C. F. de Winter, Dimitra Dodou, Harleigh C. Seyffert, Yke Bauke Eisma

arXiv 2610.05463首次发表:更新:

AI 中文总结

本研究通过心理物理学实验比较人类与多模态大语言模型在视觉搜索中的表现,发现MLLMs虽能复现人类的高层性能特征(错误率相关性达0.82),但在目标缺失时的响应机制存在显著差异。

AI 中文摘要

视觉搜索是一种基本的认知能力。本研究探讨多模态大语言模型(MLLMs)在视觉搜索任务中是否表现出类人的难度特征。我们使用相同的2D和3D刺激,在不同集合规模下,比较了人类(n = 1,250)和MLLMs的搜索性能。两组在特征搜索中均表现出高效性能,尤其在目标具有独特颜色时最为明显,但在联合搜索中,随着集合规模增加,性能出现下降。此外,我们发现人类与MLLM的错误率之间存在强相关性($\ ho = 0.82$),这表明MLLMs对类似的目标复杂度(如刺激异质性)敏感。然而,也存在差异:人类在目标缺失试验中会投入额外的搜索时间以准确响应,而MLLMs在复杂搜索中表现出极端的呈现/缺失响应偏差。我们得出结论:MLLMs复制了人类的高层性能特征,但其底层计算存在显著差异。

英文摘要

Visual search is a fundamental cognitive ability. This study investigates whether Multimodal Large Language Models (MLLMs) exhibit human-like difficulty signatures in visual search tasks. We compared search performance of humans (n = 1,250) and MLLMs using identical 2D and 3D stimuli across different set sizes. Both groups showed efficient performance in feature searches, most clearly when the target had a unique color, but performance degradation in conjunction searches as set sizes increased. Additionally, we found strong correlations between human and MLLM error rates ($ρ= 0.82$), which suggests that MLLMs are sensitive to similar objective complexities, such as stimulus heterogeneity. However, differences were found as well: whereas humans invested extra search time to respond accurately on target-absent trials, MLLMs exhibited extreme present/absent response biases in complex searches. We conclude that MLLMs replicate high-level human performance signatures, yet their underlying computations differ significantly.

Journal refComputational Brain & Behavior (2026)

DOI:10.1007/s42113-026-00333-4

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑