arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16514cs.CVcs.AIcs.CLcs.HCcs.MM

匹配结果,不同注视:中央凹多模态大语言模型(MLLM)的搜索方式与人类的对比

Matched Outcomes, Divergent Gaze: How Foveated MLLMs Search Compared to Humans

  • F-initiatives(F计划)
  • Université Sorbonne Paris Nord(巴黎北索邦大学)
  • Northwestern University(西北大学)
  • IULM university(IULM大学)

机构由 AI 辅助整理,请以论文原文为准。

Mohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani, Ulas Bagci, Alessandro Bruno

中文总结 AI 辅助

研究对比三款通用MLLM与人类在目标导向视觉搜索中的表现,发现模型在决策和目标获取上优于人类,但注视过程与人类不同,现有指标无法验证类人视觉,零样本模型不适用于过程层面问题。

中文摘要 AI 辅助

人类的视觉搜索是串行的:中央凹必须落在候选目标上才能确认,这些落点形成了扫描路径。给定相同的中央凹输入时,多模态大语言模型(MLLM)是否会像人类一样搜索,这关系到它们能否作为人类视觉的模型,以及注意力对齐分数的有效性。我们在目标导向搜索任务(COCO-Search18)上,将三款通用MLLM的眼动扫描路径与人类的进行对比,通过逐帧为每个模型提供与人类匹配的中央凹视图,从三个维度对其进行评估:目标存在与否的决策、到达目标的效率,以及注视过程本身。这三个维度是相互分离的。在决策和目标获取方面,模型的表现与人类相当或超过人类,检测存在的目标的准确率接近天花板,且首次扫视到达目标的频率高于人类。但注视过程并非人类式的。在人类匹配条件下,三款模型都有一个共同特征:低熵、大振幅、自洽的扫描路径,其一致性远高于两个人类之间的一致性。这与单次通过、非串行的架构一致,而非视敏度的限制。匹配的视网膜输入能复现人类注视的位置,但无法复现注视随时间的展开方式,且没有任何退化机制能在人类水平的成功率下恢复类人搜索。这种差异存在于过程维度,是答案对齐和显著性指标无法测量的。由于这些指标忽略了这一点,它们无法验证类人视觉,而零样本模型适用于结果和空间问题,但不适用于时间、过程层面的问题。

英文摘要

Human visual search is serial: the fovea must land on a candidate to confirm it, and those landings form a scanpath. Whether multimodal large language models (MLLMs), given the same foveated input, search as humans do bears on their use as models of human vision and on attention-alignment scores. We compare three general-purpose MLLMs with human eye-movement scanpaths on goal-directed search (COCO-Search18), driving each model fixation by fixation through an identical, human-matched foveated view and assessing it along three axes: the decision of target presence, the efficiency of reaching the target, and the gaze process itself. The axes dissociate. On the decision and on target acquisition the models match or exceed humans, detecting present targets near ceiling and reaching them on the first saccade more often than people do. The gaze process is not human. Under the human-matched condition, all three share one signature: low-entropy, large-amplitude, self-consistent scanpaths that agree with themselves far more closely than two humans agree with each other. That is consistent with a single-pass, non-serial architecture rather than a limit of acuity. Matched retinal input reproduces where humans look but not how the looking unfolds in time, and no degradation regime recovers human-like search at human-like success. The gap sits on a process axis that answer-alignment and saliency metrics do not measure. Because they miss it, such metrics cannot certify human-like vision, and zero-shot models suit outcome and spatial questions but not temporal, process-level ones.

补充信息

↑