arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

错觉模式感知导致大型语言模型中的虚假推断

Illusory Pattern Perception Drives Spurious Inference in Large Language Models

Peihua Mai, Zhuoyan Shao, Xinbao Qiao, Meng Zhang, Xinyue Zhou, Yan Pang

arXiv 2610.07791首次发表:更新:

发表机构

National University of Singapore (Chongqing) Research Institute; National University of Singapore; Zhejiang University; The Chinese University of Hong Kong; University of Macau(新加坡国立大学重庆研究院; 新加坡国立大学; 浙江大学; 香港中文大学; 澳门大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究首次系统探索LLMs中的错觉模式感知,发现其比人类更易从随机数据推断关系,并开发基于SAEs的可解释性框架揭示机制,影响推理可靠性。

AI 中文摘要

错觉模式感知是人类一种有据可查的认知倾向,即从实际随机数据中推断出有意义的关系。这种倾向常被描述为“连接不存在点”,可能导致系统性推理错误。本文研究大型语言模型(LLMs)是否表现出此类感知倾向,从而在下游应用中导致系统性错误。据我们所知,本研究首次系统性地探讨了LLMs中的错觉模式感知,将经典心理学范式应用于三项任务,并与人类行为进行直接实证比较。我们发现,LLMs经常表现出比人类更强的错觉模式感知。特别是,模型倾向于将频繁出现的正面属性与多数群体或大型组织过度关联,并表现出从模糊事件构建因果叙述的增强倾向。为揭示这些行为背后的机制,我们开发了一个基于稀疏自编码器(SAEs)的特征可解释性框架来分析内部表征。我们的结果显示,整体频率感知和分析性认知取向与错觉感知的出现相关。这些发现突显了一种先前未被充分探索的类似认知的错觉,可能影响LLM推理的可靠性。代码可在https://this URL获取。

英文摘要

Illusory pattern perception is a well-documented human cognitive tendency to infer meaningful relationships in data that is actually random. Such a tendency, often described as "connecting the dots" where none exist, can result in systematic reasoning errors. This paper investigates whether Large Language Models (LLMs) exhibit such perceptual tendencies, which can lead to systematic errors in downstream applications. To our knowledge, this work presents the first systematic study of illusory pattern perception in LLMs, adapting classic psychological paradigms to three tasks with direct empirical comparison to human behaviors. We find that LLMs frequently exhibit stronger illusory pattern perception than humans. In particular, models tend to over-associate frequent positive attributes with majority groups or large organizations, and show increased tendencies to construct causal narratives from ambiguous events. To uncover the mechanism behind these behaviors, we develop a feature interpretability framework based on Sparse Autoencoders (SAEs) to analyze internal representations. Our results reveal that holistic frequency perception and analytic cognitive orientation are linked to the emergence of illusory perceptions. These findings highlight a previously underexplored cognitive-like illusion that may affect the reliability of LLM reasoning. Code available at https://github.com/NusIoraPrivacy/illusory.

Commentsaccepted by NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑