arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型视觉语言模型(LVLMs)能否揭示视觉错觉背后的真相?感知与推理能力分析

Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities

Liangjie Zhao, Jiaqing Lyu, Kexin Tang, Zecheng Fang, Rong Yin, Yulan Hu, Da Li, Jianing Li

arXiv 2607.27747首次发表:更新:

发表机构

Adelaide University; Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences; Tsinghua University; Amap, Alibaba Group; Beihang University(阿德莱德大学; 中国科学院计算技术研究所; 中国科学院大学; 清华大学; 阿里巴巴集团高德地图; 北京航空航天大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文利用IllusionReasoning基准,以视觉错觉为诊断工具评估LVLMs,发现多款LVLMs的推理能力不及宣称水平,为其优化提供了新方向。

AI 中文摘要

大型视觉语言模型(LVLMs)已整合推理能力,将认知表现提升至新水平。然而,现有评估要么仅聚焦感知,要么依赖数学、编码等特定领域,仍需面向开放世界环境的推理能力评估,尤其需同时考量感知与推理。为填补这一空白,本文提出利用视觉错觉作为诊断工具评估LVLMs:视觉错觉是人类视觉系统误读客观信号、产生与现实相悖理解的现象。本文构建了IllusionReasoning基准,该基准包含从现实世界收集的错觉图像及多样的标注问答对。基于IllusionReasoning,研究表明多款LVLMs的推理能力并不如宣称的那般先进。本研究为LVLMs提供了新见解,并为其优化指明了未来方向。

英文摘要

Large Vision Language Models have integrated reasoning capabilities, elevating cognitive performance to new levels. However, existing evaluations either focus solely on perception or rely on specific domains such as maths or coding. Evaluation for reasoning capabilities that align with an open-world environment is still required, especially one that considers perception and reasoning jointly. To bridge this gap, we propose to evaluate LVLMs by exploiting visual illusions as a diagnostic tool. Visual illusions are phenomena in which the human visual system misinterprets objective signals, resulting in an understanding that deviates from reality. We constructed Illusion-Reasoning, a benchmark of illusion images collected from the real world, incorporating diverse annotated question-answer pairs. Based on Illusion-Reasoning, we show that the reasoning capabilities of a wide range of LVLMs are not as advanced as claimed. Our work provides new insights into LVLMs and offers future directions for optimisation. Our project is publicly available at https://github.com/zhaoliangjie55/EMNLP2026_Illusion.

CommentsEMNLP2026 Findings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑